Skip to content

An independent study reference written by Dr Phuc V. Nguyen. It is not official subject material — for assessment requirements always follow your subject outline and vUWS.

Experiments and A/B testing

An A/B test is a controlled experiment run on live users. People are assigned at random to version A or version B, and one metric chosen in advance is compared between the groups. Randomisation is what licenses the causal reading, because it balances the two groups in expectation on everything except the change, including traits nobody thought to record. It does not make the realised groups alike, so any two arms still differ by chance, and that leftover variation is what the statistics are there to measure. The hard parts are not the arithmetic. They are choosing the metric before you look, running long enough to detect a difference worth acting on, and resisting the pull to stop the moment the numbers look good.

Try it yourself

One experiment, two arms

Visitors are split at random between a control page and a treatment page. Set the true difference the simulation should build in, set how many users each arm gets, then read what one run reports. The difference is shown in percentage points first, with the relative figure worked out from it and the control rate, because a relative lift on a small base flatters a tiny change.

Observed difference
+0.58 pp
Relative lift
+14.4%
Verdict
Above zero, bar unresolved
Control1,003 of 25,0004.01%Treatment1,147 of 25,0004.59%

Conversion rate per arm, bar track fixed from 0% to 30%.

Difference in percentage points, treatment minus control0.000.250.500.751.00no changeworthwhile bar+0.58 pp
z (pooled null) 3.17p (two-sided) 0.00295% interval +0.22 pp to +0.93 pp
One realised experiment. Every statistic reported for this run is computed from the counts in this table, so a recomputation that starts from a rounded percentage instead can move the last digit.
ArmUsersConversionsRate
Control25,0001,0034.01%
Treatment25,0001,1474.59%
Baseline conversion rate4.0%
True difference built in+0.5 pp
Users per arm25,000
Smallest worthwhile effect+0.30 pp
Chance draw 1. The draw moves only when you press Run again, so a slider re-prices the same luck rather than reshuffling it.

Standard errors here: 0.181 pp pooled under the null, which is what the z-test divides by, and 0.181 pp unpooled, which is what the interval uses. The first assumes the two arms share one rate, the second does not, so they answer different questions and neither stands in for the other.

Detecting a true +0.5 pp on a 4.0% base needs roughly 24,576 users per arm from the rule of thumb on this page, which reads the variance at the baseline rate alone, against the 25,000 set here. Both of the sample sizes here are for a two-sided test at the 5 per cent level with 80 per cent power.

Working the same design from both planned rates, 4.0% against 4.5%, gives 25,551 users per arm. The two land within a tenth of each other here, because the two rates are close enough for one variance to stand in for both. The rule of thumb is a near-baseline approximation, not a general calculator.
Control converted 1,003 of 25,000 users, 4.01%, and treatment converted 1,147 of 25,000, 4.59%. That is a difference of +0.58 pp, which is a relative change of +14.4% against the control rate. The pooled-null z is 3.17 and the two-sided p-value is 0.002. The interval, +0.22 pp to +0.93 pp, clears zero but straddles the +0.30 pp bar you set. On the interval the two versions are separated, and whether the gain is worth the cost of shipping is still open. A p-value ranks how surprising this data would be if the two versions performed identically. It is not the probability that they perform identically, and it carries no information about whether the change pays for itself, so it is never on its own a decision to ship.
Simulated data. Conversions are drawn from an exact binomial distribution at the rates the sliders set, so the true difference is known here and is never known in a real test.

Why it matters

Two identical shops on the same street, the same passing crowd, the same weather. Change the window display in one and count what happens. If the display is the only systematic difference, what changed in sales is attributable to it. Assigning visitors at random online does the same job, and does it better, because chance splits the crowd on traits you could never have measured. It splits them evenly on average rather than exactly, which is why a small gap between the two arms still has to be weighed against ordinary variation.

Before you read on — recall

A test on 900 visitors per arm shows conversion of 4.0 per cent in control and 4.6 per cent in treatment. The team reports a 15 per cent lift. The best response is

Formulas

Absolute and relative lift
Δ=pT−pC,lift=pT−pCpC\Delta = p_T - p_C, \qquad \text{lift} = \frac{p_T - p_C}{p_C}
Here pCp_C is the conversion rate in the control group and pTp_T the rate in the treatment group. Moving a rate from four per cent to four and a half is an absolute gain of half a percentage point and a relative lift of about twelve and a half per cent. Reporting only the relative figure flatters small changes on small bases, so show both.
Standard error of a difference in two rates
SE=pC(1−pC)nC+pT(1−pT)nTSE = \sqrt{\frac{p_C(1-p_C)}{n_C} + \frac{p_T(1-p_T)}{n_T}}
The yardstick the observed gap is measured against. With rates near four per cent and 25,000 users in each arm, SESE is roughly 0.18 percentage points, so a half-point gap is close to three standard errors. A gap that large would be unlikely if the two versions performed identically, which is what a small p-value reports. With 1,000 users per arm, SESE is about 0.9 points and the very same gap tells you nothing at all.
How many users each group needs
n≈16 σ2Δ2per groupn \approx \frac{16\,\sigma^2}{\Delta^2} \quad \text{per group}
A standard rule of thumb for detecting a difference of size Δ\Delta under the usual conventions. For a conversion rate, σ2=p(1−p)\sigma^2 = p(1-p). Detecting a lift from four per cent to four and a half needs about n≈16×0.0384/0.000025n \approx 16 \times 0.0384 / 0.000025, which is roughly 24,600 users in each group. Small effects need large samples, and that arithmetic belongs before the test rather than after it.

Worked examples

Scenario

An online retailer tests a new checkout button. After two days the treatment is ahead by a wide margin and the product manager wants to ship it.

Solution

Two days is usually too early for two separate reasons. The sample may still be small enough that the gap sits inside ordinary random variation. The larger problem is peeking. If you check repeatedly and stop the first time the result looks good, you have stacked the deck, because noise crosses the threshold sooner or later even when the two versions are identical. Fix the run length and the metric in advance, or use a method designed for continuous monitoring. Shopping behaviour also varies from one day to the next, so a two-day test misses part of the population.

Scenario

A subscription business tests a redesigned sign-up page. Sign-ups rise nine per cent and the team declares victory.

Solution

Ask what happened after sign-up. A page that promises more than the product delivers lifts the metric being watched and raises cancellations two months later. Pair the primary metric with guardrails, here retention at 60 days and support contacts per new customer. Then check the split itself. A treatment share of 52 per cent looks close to half, but across 40,000 eligible visitors it sits about eight standard errors away from the configured split, which is a sample ratio mismatch and would almost never arise by chance. That points at the assignment or the logging rather than at the design, and until it is explained the result cannot be trusted however large the effect looks. On a few hundred visitors the same 52 per cent would be unremarkable, so the size of the sample decides whether the split is evidence of anything.

Common mistakes

  • ✗A statistically significant result is a result worth shipping. Significance says that a difference at least this large would be unlikely if the two versions performed identically. It is a statement about the data given that assumption, not the probability that the assumption is false, and it says nothing about whether the difference is large enough to pay for the change, so decide the smallest worthwhile effect before the test runs.
  • ✗You can stop a test as soon as it reaches significance. Checking repeatedly and stopping on a good look inflates false positives badly, because random noise wanders across the threshold on its own. Fix the sample size in advance, or use a sequential method built for monitoring.
  • ✗Randomisation is a formality that could be replaced by matching on known traits. Matching balances only what you thought to measure. Random assignment balances the unmeasured traits as well, in expectation rather than exactly, so whatever imbalance remains is chance rather than something systematic. That is the reason the comparison supports a causal claim, and the reason the claim comes with an interval rather than a certainty.
  • ✗A test showing no difference was a waste of traffic. A credible null result tells you not to spend on the change and rules out a story the team believed. That is worth what the test cost.

Revision bullets

  • •Random assignment balances groups in expectation, which licenses the causal reading
  • •Choose the metric and the run length before the test starts
  • •Report absolute and relative change, since relative alone flatters small effects
  • •Sample size grows with variance and with the inverse square of the effect
  • •Peeking and stopping early inflates false positives
  • •Pair the primary metric with guardrails that catch damage elsewhere

Quick check

A test on 900 visitors per arm shows conversion of 4.0 per cent in control and 4.6 per cent in treatment. The team reports a 15 per cent lift. The best response is

Which finding most undermines the validity of a completed A/B test?

Connected topics

More in Decisions and Models

Sources

  1. Kohavi, Tang & Xu (2020)
    Kohavi, R., Tang, D., & Xu, Y. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press, 2020.
    Practical treatment of online experiments, covering sample-size arithmetic, guardrail metrics, mismatched traffic splits and the cost of stopping early.
  2. Kohavi, R., Longbotham, R., Sommerfield, D., & Henne, R. M. "Controlled experiments on the web: survey and practical guide." Data Mining and Knowledge Discovery, 18(1), 2009.
    Survey setting out the mechanics of web experimentation and the recurring ways results get misread.
  3. Fisher (1935)
    Fisher, R. A. The Design of Experiments. Oliver & Boyd, 1935.
    Origin of randomisation as the device that makes treatment groups comparable on unmeasured as well as measured traits.
How to cite this page
Dr. Phil's Quant Lab. (2026). Experiments and A/B testing. Business Analytics Atlas. https://phucnguyenvan.com/analytics_atlas/concept/ba-experiments-ab-testing
Next concept
Correlation and causation
Built by Dr. Phuc V. Nguyen ·Follow on LinkedInWork with PhilEmail