Variant A (control)
Variant B (the challenger)
Free A/B test significance calculator
This free tool — built by a Singapore digital marketing agency — tells you whether the difference between two versions of a page, email or ad is a real improvement or just random noise. Enter the visitors and conversions for your control (A) and your challenger (B), and it runs a two-proportion z-test to return the conversion rates, the relative uplift, and the statistical confidence that B genuinely differs from A. It runs entirely in your browser — nothing is uploaded.
What “statistical significance” actually means
Every test has random variation: flip a fair coin 100 times and you rarely get exactly 50 heads. Statistical significance answers a simple question — if A and B were really identical, how often would we see a gap this big by pure chance? A confidence of 95% means there’s only a 5% probability the result is a fluke. That 95% threshold is the widely used convention for calling a test; below it, you risk shipping a “winner” that was really just luck. This calculator reports the two-tailed confidence so you can see exactly where you stand.
How many visitors do you need?
The smaller the true difference between A and B, the more traffic you need to prove it. A jump from 5% to 10% shows up fast; a 5% to 5.3% lift can take tens of thousands of visitors per variant. As a rough guide, treat any test with fewer than ~1,000 visitors and ~100 conversions per variant as underpowered — interesting, but not decisive. Resist the urge to “peek” and stop the moment it crosses 95%; decide your sample size in advance and let the test run its course to avoid false positives.
Common A/B testing mistakes we see in Singapore
The most frequent errors: calling a test after two days because one variant looks ahead; running tests during an atypical week (a sale, a public holiday, a viral spike); testing tiny changes that could never move the needle enough to detect; and ignoring organic and AI-search traffic segments that behave very differently from paid visitors. A significant result on a representative sample, run for at least one full business cycle, is worth far more than a fast one.
Frequently asked questions
What confidence level should I use?
95% is the standard for most marketing tests — a good balance between not shipping flukes and not waiting forever. High-stakes changes (checkout, pricing) may warrant 99%; low-risk copy tweaks can act on 90% with more monitoring. This tool flags 95%+ as a call, 90–95% as a lean, and below 90% as inconclusive.
What is relative uplift?
It’s the percentage improvement of B over A. Going from a 5% to a 6% conversion rate is a 1 percentage-point absolute gain but a 20% relative uplift ((6−5)÷5). Relative uplift is the figure most teams quote because it scales with your baseline.
Is this a one-tailed or two-tailed test?
Two-tailed — it tests whether B is different from A in either direction, which is the safer default. That’s why a variant that is clearly worse also reaches high confidence: you can be confident it’s a real change, just not a positive one.
Can I use this for email or ad tests?
Yes. Any two-variant test with a binary outcome works: opens vs sends, clicks vs impressions, sign-ups vs visitors, purchases vs sessions. Just enter the “attempts” as visitors and the “successes” as conversions for each variant.