How long should an A/B test run?

The honest answer: until the sample size you calculated up front is reached, but at least two full weeks.

What this article covers
  1. Two numbers determine the runtime
  2. Why at least two weeks anyway
  3. The most expensive mistake: peeking and stopping
  4. When you should stop a test anyway
  5. The rule of thumb

“We saw +18% after three days, can we ship it?”, that question comes up in almost every project. The short answer is no. The long answer explains why those +18% after three days are almost never real.

Two numbers determine the runtime

How long a test has to run does not depend on gut feeling but on two numbers: your current conversion rate and the smallest difference you still want to detect. The lower the baseline and the smaller the difference you are looking for, the more visitors you need.

An example: with a 3% baseline and a 10% relative difference you are looking for, you need roughly 50,000 visitors per variation. With a 10% baseline it is around 14,000. You do this calculation before launch, not once the numbers already look exciting.

Why at least two weeks anyway

Even if the sample is complete after four days, keep going. Buying behaviour on a Monday differs from a Saturday, and measuring only weekdays measures a slice. Two full weeks cover every weekday twice and absorb a single outlier day.

The reverse also holds: a test running longer than four to six weeks becomes fragile. Cookies expire, visitors switch devices, campaigns change the traffic mix. If you need that long, the change you are testing is probably too small.

The most expensive mistake: peeking and stopping

Checking the dashboard every day and stopping at the first green spike produces false winners systematically. The reason is simple: random swings are large early on. Look often enough and one variation eventually looks good by chance, and that is exactly when you stop.

So: fix runtime and sample size in advance, then leave them alone. You may look, you just may not decide.

When you should stop a test anyway

  • The variation is visibly broken, console errors, wrecked layout, abandoned purchases.
  • The traffic split is off, say 60/40 instead of 50/50. Then something is wrong with delivery, not with the idea.
  • A guardrail fires: the variation improves the main goal but clearly worsens returns or support requests.
  • An outside event distorts everything, an outage, press coverage, Black Friday in the middle of your test window.

The rule of thumb

Calculate the sample size in advance, run for at least two full weeks and at most six. Do not decide before that, and read honestly afterwards: “no difference” is a result too, it saves you from building an idea that does nothing.

Make your next change an informed one

Less guesswork.
More now we know.

One question is a good place to start.