How to Know If Your Ad Test Actually Won

August 29, 2026 · 7 min read · By James Leary


An agency reports that the new ad is winning. Version B converted at 8% against version A’s 5%, so B gets the budget and A gets switched off.

Then the next month the numbers move back, and nobody can explain why.

Usually nothing went wrong with the ads. The test was called on 40 clicks and 3 conversions, where an 8% versus 5% gap is well inside the range you would get by flipping a coin. The winner was noise, and switching off A cost real money.

Why contractor tests get called early

Contractors run into this harder than most advertisers, for a structural reason: the conversions are rare and expensive.

An ecommerce store testing a checkout button gets thousands of conversions a week and can settle a question by Thursday. A roofing company generating 40 leads a month gets 40 data points. At that volume, ordinary week-to-week variation swamps most differences you would care about.

The instinct is to call it anyway, because waiting feels like wasting spend on a losing variant. But calling it early does not save money — it moves the loss somewhere you cannot see, into a “winner” that was never actually better.

What significance actually means

Statistical significance answers one narrow question: if the two versions were genuinely identical, how often would I see a gap this big by chance alone?

That is all. A p-value of 0.05 means a gap this large would turn up about one time in twenty even if both ads were equally good. It does not mean B is 95% likely to be better, and it does not tell you the difference is large enough to matter commercially.

The A/B test significance calculator runs the standard two-proportion z-test on your own numbers: visitors and conversions for each version, out come the conversion rates, the relative lift, the p-value and a confidence level.

If it tells you the result is not significant, that is not a failed test. It is the test telling you it does not have enough data to answer yet.

The sample size problem, honestly

Here is the uncomfortable arithmetic. Detecting a genuine difference reliably needs more data the smaller that difference is — and small differences need a lot.

Going from a 5% conversion rate to 10% is a big, obvious change and shows up quickly. Going from 5% to 6% is a 20% improvement in your business, genuinely worth having, and needs thousands of visitors per variant to separate from noise.

For most contractors that means:

  • Big swings are testable. A completely different offer, a different landing page structure, a different audience. If it moves the number by half again, you will see it.
  • Small refinements usually are not. Button colour, a reworded headline, a slightly different image. The effect may well be real and you will never prove it at your volume.

That is not a reason to stop testing. It is a reason to test things big enough to see, and to stop pretending the small stuff was measured.

Test one thing, on a metric that matters

Two rules separate a test worth running from theatre.

Change one thing. If B has a new headline, a new image and a new form, a win tells you nothing about which change caused it — and you cannot carry the lesson to the next campaign.

Measure at the bottom, not the top. Optimising for the cheapest cost per lead reliably produces cheaper, worse leads. Judge variants on cost per booked job where you have the volume, and on qualified-appointment rate where you do not. The argument for that metric is worth reading before setting up any test.

Decide the rules before you look

The most common way tests go wrong is not statistical, it is human: watching the numbers daily and stopping the moment they look good. Check a running test often enough and it will cross any threshold you like at some point, purely by chance.

Fix it by writing down three things before the test starts:

  1. The metric. One, chosen in advance.
  2. The stopping point. A number of conversions per variant, or a fixed run length — at least one full business cycle, so a test never spans a holiday week on one side and a normal week on the other.
  3. The size worth acting on. If a 3% lift would not change what you do, do not run a test that can only detect 3%.

Then leave it alone until the stopping point, and run the numbers once.

When you cannot reach significance, decide anyway

Most contractor tests will end inconclusive. That is the honest outcome at low volume, and it does not leave you stuck.

If the result is not significant, the two versions have not been shown to differ — so choose on other grounds. Which is easier to maintain? Which reflects the offer you actually want to sell? Which produced leads your crew preferred? Those are legitimate reasons. What is not legitimate is reporting the inconclusive result as a win, which is how a three-conversion difference becomes a permanent strategy nobody revisits.

If an agency reports winners without sample sizes or a p-value, that is worth asking about — it belongs with the other questions in how to tell if your marketing agency is working.

Frequently Asked Questions

How many conversions do I need before an A/B test is meaningful?

It depends on the size of the difference you are trying to detect, not on a fixed number. A large difference — one version converting half again as well as the other — can show up in a few dozen conversions per variant. Detecting a 20% relative improvement typically needs hundreds to thousands of visitors per variant. Run your own numbers through a two-proportion z-test rather than trusting a rule of thumb, and treat “not significant” as the test saying it has not seen enough data yet.

What does statistical significance actually tell me?

Only this: if the two versions were genuinely identical, how often would a gap this large appear by chance. A p-value of 0.05 means about one time in twenty. It does not mean the winner is 95% likely to be better, and it says nothing about whether the difference is big enough to matter commercially. Those remain separate judgements you have to make yourself.

Why do my ad tests keep reversing the next month?

Almost always because they were called too early. At contractor lead volumes, ordinary week-to-week variation is large relative to the differences being tested, so a version can look clearly ahead on 40 clicks and mean nothing. Deciding the stopping point before the test starts, and not checking daily, prevents most of this.

Can small contractors run A/B tests at all?

Yes, but only on changes big enough to see. A different offer, a different landing page structure or a different audience can move conversion rates enough to detect at modest volume. Small refinements like button colour or a reworded headline usually cannot be separated from noise at contractor volumes, and pretending otherwise produces confident conclusions built on nothing.

What should I measure in an ad test?

The metric closest to money that you have enough volume to measure. Optimising for the cheapest cost per lead reliably produces cheaper, worse leads, so judge variants on cost per booked job where volume allows and on qualified-appointment rate where it does not. Whichever you pick, choose it before the test starts rather than after seeing which metric favours the version you already liked.


Testing is worth doing. Calling a winner on three conversions is not testing — it is deciding, then dressing it up as measurement afterwards.

See what our campaigns produce, or book a call and we will look at what your volume can actually measure.

25 client cap

Want this handled for you?

20 qualified appointments in 30 days, guaranteed in writing. Twenty minutes tells you whether your market can support it.

Book a Free Call →