A/B test significance calculator
An A/B test is significant when the gap between two conversion rates is too large to blame on chance. Version A: 120/4,000 = 3.00%. Version B: 156/4,000 = 3.90%. That is a +30% relative lift at 97.3% confidence — a real winner. Below 95%, keep the test running.
Enter both versions
Everyone who saw the control.
Form fills, calls, whatever you count as the win.
Everyone who saw the variant.
Counted exactly the same way as version A.
Version B wins at 97.3% confidence. A gap this large would appear by chance about 2.7 times in 100.
95% or above is the conventional bar for calling a winner.
An absolute gap of 0.90 percentage points.
Probability of a gap this large if the two versions were identical.
Two-proportion z-test on the pooled rate. Roughly 1.96 is the 95% line.
Both versions have enough conversions and non-conversions for the test to be valid.
How to use this calculator
Enter how many people saw each version and how many of them converted. Everything else is derived, and the results update as you type. Count conversions identically on both sides — if a phone call counts as a conversion on version A, it has to count on version B, and both versions have to have run over the same days.
Read the verdict first, then the sample check. A significant verdict on an inadequate sample is not a finding, it is noise that happened to line up, and the calculator will say so rather than letting you act on it.
A worked example
A roofing company tests two landing pages. The control gets 4,000 visitors and 120 form fills, a conversion rate of 3.00%. The variant — same traffic, form cut from eight fields to three — gets 4,000 visitors and 156 form fills, or 3.90%.
The absolute gap is 0.90 percentage points. The relative lift is +30%. Pooling the two rates gives a z-score of 2.21 and a two-tailed p-value of 0.027, so the confidence level is 97.3%. That clears the 95% bar and the shorter form ships.
Now run the same test at a tenth of the traffic: 400 visitors and 12 conversions against 400 and 16. The rates are identical, the lift is still 30%, and confidence drops to roughly 55% — no better than a coin toss. The improvement did not change. The evidence did.
Reading your own result
- Confidence at or above 95%. Ship the winner. Note the size of the lift, not just the fact of it, because a significant 2% lift may not be worth the work of implementing it.
- Confidence between 80% and 95%. Suggestive, not conclusive. Keep the test running if traffic allows. Do not report it as a win.
- Confidence below 80%. No evidence either way. Either the change was too small to detect or the sample is too thin. Test something bolder rather than waiting this one out.
- Sample check flags a problem. One version has too few conversions or too few non-conversions for the normal approximation behind the z-test to hold. Whatever the confidence number says, do not act on it.
The common mistake
Watching the test daily and stopping the moment it crosses 95%. A result that wanders will cross that line by chance at some point in almost any test, so stopping on the crossing turns a one-in-twenty false positive rate into something far worse. Decide the sample size and the end date before you start, then look once. If a result was significant on day three and gone by day nine, it was never there.
The second mistake is testing changes too small to matter. Most local contractors do not have the traffic to resolve a button colour. They do have enough to resolve a different offer, a form cut from eight fields to three, or a service-specific page against a general homepage. Test things big enough to show up.
When you have a winner, work out what it is worth with the marketing ROI calculator, check the click economics behind it with the CPC and CTR calculator, and confirm the funnel underneath can carry the extra volume with the cost per booked job calculator.
Common questions
What does statistical significance actually mean in an A/B test?
Statistical significance means the gap between two versions is large enough, given the sample size, that random chance is an unlikely explanation for it. A result significant at 95% confidence would appear by luck alone roughly one time in twenty when the two versions are genuinely identical. It is a statement about how much the evidence can be trusted, not a statement about how much money the winner will make.
What is a two-proportion z-test?
It is the standard test for comparing two conversion rates. The two rates are pooled to estimate what the shared rate would be if the versions were identical, that pooled rate gives the standard error of the difference, and the observed difference is divided by that standard error to produce a z-score. The z-score converts to a p-value, which is the probability of seeing a gap this large by chance.
What is a p-value?
The p-value is the probability of observing a difference at least as large as the one measured, assuming the two versions actually perform identically. A p-value of 0.03 means a gap this size would turn up by chance about three times in a hundred. Lower is stronger evidence. The conventional cut-off for calling a winner is 0.05, which corresponds to 95% confidence.
How much traffic does an A/B test need?
It depends on the baseline conversion rate and the size of the improvement worth detecting. Low conversion rates and small improvements need very large samples, which is why a page converting at two percent may need tens of thousands of visitors per version to reliably detect a ten percent relative lift. A practical floor is at least a few hundred conversions per version before the result deserves any weight.
Why should you not stop a test as soon as it looks significant?
Checking repeatedly and stopping the moment the threshold is crossed inflates the false positive rate substantially, because a wandering result will cross the line by chance at some point during almost any test. Decide the sample size and the end date before starting, then look once at the end. A result that was significant on day three and gone by day nine was never real.
What is relative lift versus absolute difference?
Absolute difference is the gap in percentage points: 3.0% against 3.9% is 0.9 points. Relative lift expresses that gap as a proportion of the baseline: 0.9 divided by 3.0 is a 30% lift. Relative lift is the more useful number for forecasting revenue, and the more misleading one in a headline, because a 50% lift on a rate of 0.2% is still almost nothing.
Does a non-significant result mean the two versions are the same?
No. It means the test did not gather enough evidence to distinguish them. A genuinely better version can easily fail to reach significance on a small sample, and calling that a tie throws away a real improvement. The correct reading of a non-significant result is that the question is still open, and the options are to keep running, test a bolder change, or move on.
What should contractors actually A/B test?
Test things large enough to move a number: an entirely different offer, a form that asks two questions instead of eight, a call button above the fold against a form, or a landing page built for one service against a general homepage. Button colours and headline word swaps produce differences too small to detect at the traffic volumes most local contractors have.
Keep going
Cost per click, click-through rate and CPM together, from spend, clicks and impressions.
Free tool Marketing ROI CalculatorWhat the winning version is actually worth: ROI, net profit and payback multiple.
Free tool Cost Per Booked Job CalculatorA better conversion rate only matters if the funnel behind it holds. See where it leaks.
Testing headlines?
Test the offer instead.
20 qualified appointments in 30 days, guaranteed in writing. That is a result you do not need a z-test to read.
Book a Free Call →