HomeGuidesSEO & ContentThe Complete Guide to A/b Testing
SEO & Content

The Complete Guide to A/b Testing

A/B testing means showing two versions of a page, an email, or an ad to separate groups of visitors and measuring which one performs better. It sounds simple, and running one badly is even simpler: too little traffic, tests stopped the moment they look good, five changes bundled into one variant so nobody can tell which one actually mattered, or a metric that doesn’t reflect the real goal. Done right, A/B testing replaces opinion with evidence. Done wrong, it produces a confident-sounding number that means nothing, and a lot of teams ship changes based on exactly that.

Key takeaways

  • A/B testing only works with enough traffic and enough time. Testing a low-traffic page for three days produces a result that looks meaningful and usually isn't.
  • Test one variable at a time, or you won't know which change actually caused the result.
  • Statistical significance matters. A test that hasn't reached 95% confidence is a guess with a percentage attached.
  • Most tests don't win. That's normal, not a sign the process is broken.
  • The biggest wins usually come from testing the offer or the headline, not the button color.

What A/B testing actually is

A/B testing splits your visitors into two groups, shows each group a different version of the same page or element, and measures which version produces more of whatever you’re tracking: signups, purchases, clicks, whatever the goal is. Version A is usually the current page (the control). Version B is the change you’re testing (the variant).

The test runs until you have enough data to say, with real statistical confidence, that one version actually outperforms the other, rather than just looking like it did by chance. That last part is where most homegrown testing goes wrong.

The alternative, shipping a redesign because it feels right and hoping conversion improves, isn’t wrong exactly, but it removes the one thing testing provides: proof, one way or the other, instead of a guess dressed up as a decision.

What to test first

Not every element is worth testing. Button color and font choices get tested constantly and rarely move the needle enough to matter. The tests that actually produce meaningful lifts tend to hit bigger levers: the headline, the core offer, the number and order of form fields, the price or the way it’s framed, and the primary call to action’s wording.

If a page isn’t converting at all, a small test is the wrong move. Fix the obvious problems first, a confusing layout, a slow load time, a broken form, or an unclear offer, and save testing for refining something that’s already working reasonably well.

A headline or offer change can move conversion by double digits on a page that wasn’t working. A button color change rarely moves it more than a fraction of a percentage point, if it moves at all. Spend testing time where the potential lift actually justifies the setup.

How much traffic you actually need

This is where most small sites get A/B testing wrong. A valid test needs enough visitors in each variant to reach statistical significance, and for most conversion rates that means hundreds of conversions per variant, not visits. A page converting at 3% with 500 monthly visitors produces about 15 conversions a month. Split that across two variants and you’re waiting months for a result that might still be noise.

A rough industry rule of thumb: don’t start A/B testing until a page or flow gets at least a few thousand visitors a month, and even then expect tests to run for two to four weeks minimum. Lower-traffic sites are usually better off with qualitative feedback, like user interviews, than a formal test.

Minimum detectable effect matters here too. Catching a 2% lift needs a much bigger sample than catching a 20% lift, since small effects are harder to distinguish from noise. Most sample-size calculators built into testing tools will do this math for you, so there’s no need to guess.

Setting up a valid test

Decide the single metric that defines a win before the test starts, not after you’ve seen the numbers. Change one variable between control and variant. Run both versions at the same time, not sequentially, since traffic quality, seasonality, day-of-week patterns, and outside promotions all shift results if you test A this week and B the next.

Pick a sample size and a minimum run time in advance, based on your current conversion rate and traffic, and commit to it. Stopping a test early because one version is temporarily ahead is the single most common way a test produces a false result.

Consider a holdout group for bigger changes, a small slice of traffic that never sees any variant, so you can measure the test’s effect against a true baseline instead of just comparing two variants against each other.

Reading the results without fooling yourself

A version that’s ahead after two days usually isn’t actually winning. Early data is noisy, and small sample sizes swing wildly before settling down. Wait for statistical significance, typically 95% confidence, and for the sample size you calculated up front.

Watch for a segment trap too: a variant that wins overall but loses badly on mobile, or wins with new visitors and loses with returning ones. Segment the results before declaring a winner, since an average that hides a split result isn’t actually telling you what happened.

A statistically significant result isn’t automatically a practically significant one. A 0.3% lift might clear the significance bar on a huge sample and still be too small to justify the engineering time it took to build the variant. Weigh the size of the win against the cost of shipping it, not just the p-value.

Common A/B testing mistakes

  • Testing too many things at once, so a win can’t be attributed to any single change
  • Stopping the test as soon as it looks good instead of waiting for real significance
  • Running the test on too little traffic to ever produce a reliable answer
  • Ignoring external factors, a holiday sale, a press mention, a seasonal spike, or a competitor’s outage, that skew results independent of the actual test
  • Treating a losing test as wasted effort instead of a real answer about what doesn’t work

That last one matters more than it sounds. A test that proves an idea doesn’t work saves you from shipping it permanently, which is worth something even though it doesn’t feel like a win.

Tools for running A/B tests

Google Optimize shut down in 2023, which pushed a lot of small businesses toward paid tools like VWO, Optimizely, or Convert, or toward built-in testing inside platforms like Shopify and HubSpot. For teams already running digital marketing campaigns, ad platforms like Google and Meta also run their own split tests on creative and audiences, which is a lower-effort starting point than building a full page test.

Whatever tool you pick, confirm it’s actually randomizing visitors correctly and tracking the metric you think it’s tracking before trusting a single result from it.

Client-side testing tools can cause a brief flash of the original page before the variant loads, which frustrates visitors and can itself skew results. Server-side testing avoids this but takes more engineering effort to set up. Weigh that tradeoff against how much traffic and development time you actually have.

When A/B testing isn't worth it yet

Below a certain traffic threshold, formal testing costs more in time and tooling than it returns in insight. If you’re not sure whether your site qualifies, the honest answer for most local businesses and early-stage sites is: not yet. Fix obvious usability problems, talk to five real customers, and revisit testing once traffic and conversions are high enough to produce a real answer in a reasonable timeframe.

If you want a read on whether your traffic supports meaningful A/B testing, or want help structuring the first one, reach out and we’ll look at the numbers with you.

A simpler alternative for low-traffic sites is a before-and-after comparison: make the change, watch performance over a comparable period, and adjust if it doesn’t hold up. It’s not as rigorous as a true test, since seasonality and other factors aren’t controlled for, but it beats doing nothing while waiting for traffic to grow.

Ready to build the whole thing right?

One studio, one system, from first mark to full scale.

Start a project

Frequently asked questions

How long should an A/B test run?
At minimum two to four weeks, and only after it hits the sample size you calculated in advance. Shorter A/B testing runs are usually just noise dressed up as a result.
What's a good sample size for A/B testing?
It depends on your current conversion rate and the size of the change, but most valid tests need hundreds of conversions per variant, not just visits, before the result means anything.
Can I test more than one thing at a time?
You can, with a multivariate test, but it needs far more traffic than simple A/B testing and gets harder to interpret. For most sites, one variable at a time is the practical answer.
Is A/B testing worth it for a low-traffic website?
Usually not yet. Below a few thousand monthly visitors to the page being tested, A/B testing takes too long to reach a reliable answer. User interviews and direct feedback are more useful at that stage.
What should I do if my A/B test doesn't win?
Treat it as a real answer, not a failure. A losing variant tells you what doesn’t move the needle, which keeps you from shipping that change permanently based on a hunch.
Start a project