Skip to content
One experiment, two arms

Visitors are split at random between a control page and a treatment page. Set the true difference the simulation should build in, set how many users each arm gets, then read what one run reports. The difference is shown in percentage points first, with the relative figure worked out from it and the control rate, because a relative lift on a small base flatters a tiny change.

Observed difference
+0.58 pp
Relative lift
+14.4%
Verdict
Above zero, bar unresolved
Control1,003 of 25,0004.01%Treatment1,147 of 25,0004.59%

Conversion rate per arm, bar track fixed from 0% to 30%.

Difference in percentage points, treatment minus control0.000.250.500.751.00no changeworthwhile bar+0.58 pp
z (pooled null) 3.17p (two-sided) 0.00295% interval +0.22 pp to +0.93 pp
One realised experiment. Every statistic reported for this run is computed from the counts in this table, so a recomputation that starts from a rounded percentage instead can move the last digit.
ArmUsersConversionsRate
Control25,0001,0034.01%
Treatment25,0001,1474.59%
Baseline conversion rate4.0%
True difference built in+0.5 pp
Users per arm25,000
Smallest worthwhile effect+0.30 pp
Chance draw 1. The draw moves only when you press Run again, so a slider re-prices the same luck rather than reshuffling it.

Standard errors here: 0.181 pp pooled under the null, which is what the z-test divides by, and 0.181 pp unpooled, which is what the interval uses. The first assumes the two arms share one rate, the second does not, so they answer different questions and neither stands in for the other.

Detecting a true +0.5 pp on a 4.0% base needs roughly 24,576 users per arm from the rule of thumb on this page, which reads the variance at the baseline rate alone, against the 25,000 set here. Both of the sample sizes here are for a two-sided test at the 5 per cent level with 80 per cent power.

Working the same design from both planned rates, 4.0% against 4.5%, gives 25,551 users per arm. The two land within a tenth of each other here, because the two rates are close enough for one variance to stand in for both. The rule of thumb is a near-baseline approximation, not a general calculator.
Control converted 1,003 of 25,000 users, 4.01%, and treatment converted 1,147 of 25,000, 4.59%. That is a difference of +0.58 pp, which is a relative change of +14.4% against the control rate. The pooled-null z is 3.17 and the two-sided p-value is 0.002. The interval, +0.22 pp to +0.93 pp, clears zero but straddles the +0.30 pp bar you set. On the interval the two versions are separated, and whether the gain is worth the cost of shipping is still open. A p-value ranks how surprising this data would be if the two versions performed identically. It is not the probability that they perform identically, and it carries no information about whether the change pays for itself, so it is never on its own a decision to ship.
Simulated data. Conversions are drawn from an exact binomial distribution at the rates the sliders set, so the true difference is known here and is never known in a real test.
Experiments and A/B testingOpen in Dr Phil's Quant Lab ↗