Back to Tools
A/B Test Calculator

Is Your Test Significant?

Drop in your control + variant numbers. Get conversion rates, lift, z-score, p-value, and the sample size you need for 80% power.

By Daria Dovzhikova · Updated September 2026

Inputs

Control (A)
Variant (B)
Control Conversion Rate
5%
Variant Conversion Rate
5.8%
Relative Lift
+16%
Z-Score
Standard deviations between rates
2.503
P-Value
Probability result is due to chance
0.0123
Statistical Significance
Below p=0.05 threshold
Significant @ 98.8%
Required Sample / Variant
80% power · α=0.05 · 5% MDE
119,168

What this calculator does

This tool answers two questions in one screen. First: given my control and variant conversion data, is the observed lift statistically significant or is it noise? Second: given my baseline conversion rate and the minimum effect I'd care about, how many users per arm do I need to run a properly- powered test? Both questions get the same standard frequentist treatment used by Optimizely, VWO, Google Optimize (RIP), and most in-house experimentation platforms.

The math behind it

Significance: a two-proportion z-test. The z-score measures how many standard deviations apart the two conversion rates are; the p-value is the probability of seeing that gap (or bigger) if the variants were actually identical. Convention: p < 0.05 (95% confidence) is the publish threshold. Sample size: derived from the baseline rate, the minimum detectable effect (MDE), and the desired statistical power (typically 80%) using the formula from Cohen's sample-size formula.

Common mistakes the calculator helps avoid

Peeking.If you check significance every day and stop the test the moment p < 0.05, you'll declare wins that don't exist. The published p-value assumes a single check at a pre-committed sample size. Use the sample- size estimator first, then don't look until you hit it.

Under-powered tests. A 1% baseline conversion rate and a 10% relative MDE needs roughly 31,000 users per arm for 80% power at 95% confidence. Most landing-page tests are run with 2-3K users per arm and declared inconclusive; the test was simply too small to find the effect even if it existed.

Multiple-comparison inflation. Testing five variants against a control at 95% confidence each gives you a ~22% chance of a false positive somewhere. Bonferroni-correct or pick one variant.

When to run an A/B test (and when not to)

A/B tests pay off when (a) you have enough traffic for the math to work out in under a month, (b) the change is reversible if the test fails, and (c) the cost of being wrong is meaningful. Skip the test when traffic is below a few thousand conversions per month total: you'll get a faster signal by talking to users than by running an under-powered test for 8 weeks. For early-stage product decisions, the Lean Startup validated-learning loop beats A/B testing every time.

A/B testing FAQ

How do I calculate the sample size for an A/B test? You need three numbers: the baseline conversion rate, the minimum detectable effect (the smallest relative lift worth acting on), and the statistical power you want (80% is standard). The formula this calculator uses: n = 2 x p(1-p) x (1.96 + 0.84)^2 / MDE^2 per variant, at 95% confidence and 80% power. Example: a 5% baseline with a 5% relative MDE needs roughly 120,000 visitors per arm; a 10% relative MDE needs about 30,000.

How long should an A/B test run? Divide the required sample size per variant (from this calculator) by your daily traffic per variant. A test that needs 30,000 users per arm on a page getting 2,000 visits a day split 50/50 will take 30 days. Two extra rules: run at least one full week regardless, so weekday and weekend behavior are both represented, and never stop early just because significance appeared. Stopping at the first p < 0.05 is the peeking mistake and it manufactures false winners.

What is the difference between an A/B test and a split test? Nothing. Split test, split-URL test, and A/B test describe the same method: traffic divided between a control and a variant, with conversion measured on each. Split-URL testing usually refers to variants living on different URLs rather than being injected on one page, but the statistics, and this calculator, are identical.

What does statistical significance mean in an A/B test? It means the observed difference between control and variant would be unlikely (under 5% probability, at the p < 0.05 convention) if the two versions actually performed identically. It does not mean the lift is large, durable, or worth shipping; a significant 0.2% lift can still be commercially irrelevant. Judge the effect size and the confidence interval together, not the p-value alone.

Should I use a Bayesian or frequentist A/B test calculator? This calculator is frequentist (z-test, p-values), which is the convention at Optimizely-class platforms and in most published experiments. Bayesian calculators express the result as a probability the variant beats control, which is easier to read and handles continuous monitoring more gracefully. With a pre-committed sample size and no peeking, both approaches agree on any decision worth making. The discipline matters more than the school.

Two further references

For a longer treatment of the statistics, see Evan Miller's A/B testing essays, the most-cited practitioner reference on the topic. For Bayesian alternatives to the frequentist approach this calculator uses, see Split.io's framework.

Embed this calculator

Add this to your blog, course, or internal docs. Free, no attribution removed, branded back to gtm-labs.co.

Share: LinkedIn X / Twitter

Ready when you are.

Discovery calls are 20 minutes. First one's on me.

Book a Strategy Call