Skip to content

Experiment design & evaluation

A/B Test Calculator

Plan the sample your experiment needs, then evaluate Control versus Variant results with a transparent two-proportion methodology — including Bonferroni multi-arm correction and Newcombe intervals.

Methodology posture

Built for decision-grade experiment planning

Fixed two-proportion formulas, explicit validation, and conservative withhold rules when asymptotic assumptions break down — so results stay credible for planning and evaluation.

  • Explicit Calculate and Reset actions — no silent recalculation while typing.
  • Empty inputs stay empty. Zero is never inferred from a blank field.
  • Inputs remain in memory for this page session only.

Workspace

A/B Test Calculator

Choose a mode, enter values, then calculate. Inputs stay in memory on this page only — nothing is written to the URL or storage.

Switching modes does not recalculate. Each mode keeps its own inputs and last result until you reset it or refresh the page.

Use this mode before launch to estimate the sample and duration needed to detect a meaningful conversion difference at your chosen Confidence level and Statistical power.

Load realistic example inputs to see how the calculator works.

Experiment planning inputs

Current Control conversion rate as a percentage (for example, 5 for 5%). Must be greater than 0% and less than 100%. Higher baselines usually need different sample sizes than lower ones.

Use this when uplift is described as a percentage of the current conversion rate. For example, a 20% relative uplift on a 5% baseline produces a 6% target.

Smallest relative uplift you want the experiment to detect, as a percent of the baseline. Smaller effects need larger samples. The resulting target rate must stay between 0% and 100%.

How strict the false-positive control should be. Higher confidence increases the required sample.

Chance of detecting the MDE if it is real. Higher power increases the required sample.

Count Control plus every Variant. Extra arms create more comparisons; Bonferroni correction then increases sample needs.

Whole-number estimate of daily visitors across all arms combined, not per Variant. Used only to estimate experiment duration.

Sample size is calculated only when you press Calculate. Changing inputs after a result marks that result as stale until you recalculate.

Experiment planning preview

See how your assumptions shape the test

Preview the conversion pathway from your baseline and MDE. Required sample and duration appear only after you press Calculate sample size.

Conversion pathway

Enter a baseline conversion rate and minimum detectable effect to preview the target conversion rate.

Current conversion

Minimum detectable effect

Target conversion

After calculation

  • Required sample per arm — after Calculate
  • Total required sample — after Calculate
  • Estimated duration — after Calculate

Daily experiment traffic is divided across the selected arms to estimate sample and whole-day duration.

Example: A 5% baseline with a 20% relative uplift means planning to detect an increase from 5% to 6% — not a jump to 25%.

Methodology overview

What the calculator computes

A concise overview of the statistical approach behind sample-size planning and significance evaluation.

Sample Size

Binary independent proportions with a two-proportion normal approximation. Multi-arm designs apply Bonferroni correction across arms − 1 comparisons.

Statistical Significance

Pooled two-proportion z-test, Wilson intervals per arm, and a Newcombe hybrid-score interval for the absolute difference.

Conservative decisions

Sparse expected counts, degenerate variance, or p-value/CI conflict produce a withheld decision rather than an overconfident winner.

Transparent presentation

Statistical values keep full precision internally. Rounding, percentage formatting and “Unavailable” labels apply only when results are shown.

Key assumptions

Conditions the calculations rely on

  • The primary metric is a binary conversion outcome.
  • Visitors are independent and randomly assigned across arms.
  • Conversion rates are stable enough over the planned experiment window.
  • Traffic is split evenly when estimating duration from total daily traffic.
  • Significance mode evaluates one Control vs one Variant comparison.

Practical limitations

What this tool does not claim

  • This calculator does not replace sequential testing, Bayesian decision frameworks or CUPED / covariate adjustment.
  • Sample-size estimates assume a fixed-horizon design and do not model peeking or early stopping.
  • Significance results rely on asymptotic normal approximations; very sparse counts are withheld by design.
  • Relative uplift is unavailable when Control conversion rate is zero.
  • Operational constraints — seasonality, novelty effects, instrumentation drift — are outside the statistical model.

Common mistakes

Frequent A/B testing pitfalls

Use the calculator as a planning and evaluation aid — not as a substitute for experiment design discipline.

Stopping when it “looks significant”

Repeated peeking without a sequential design inflates false positives. Size the test first, then evaluate at the planned sample.

Ignoring multiple comparisons

Extra variants increase the chance of a spurious winner. Sample Size mode applies Bonferroni correction when the experiment has more than two arms.

Oversizing MDE ambition

Tiny effects can require impractical samples. Use an MDE tied to a business decision, not the smallest imaginable lift.

Trusting colour or dashboards alone

A decision needs rates, uncertainty, assumptions and an explicit significance rule — not a green/red chart cue.

Related Growth Tools

Continue the decision path

Move from experiment maths into measurement maturity, unit economics, or the broader Growth Tools overview.

  • AI Growth Infrastructure Assessment

    Audit whether tracking, attribution and revenue data are reliable enough to trust experiment outcomes.

    Open Growth Assessment
  • Predictive LTV & Payback Calculator

    Translate conversion lifts into unit-economics impact through contribution LTV and CAC payback.

    Open LTV & Payback Calculator
  • Growth Tools overview

    Browse the live decision tools and see how assessment, modelling and experimentation fit together.

    View Tools Overview

Disclaimer

Results are provided for informational and analytical purposes only. They are based on user-provided data and statistical assumptions and do not constitute financial, legal, medical or other professional advice. Read the full Disclaimer

From the practice

Turn this experiment plan into a plan

The A/B Test Calculator gives you a disciplined sample-size and significance read. Turning experiment design into a reliable measurement and decision system is what the Lazarevych growth-analytics practice does next.

Need implementation support? Turn the result into a scoped analytics, measurement or unit-economics workstream.

Discuss Implementation with Maksym