Skip to content
Confounding and the pooled gap

A subscription business has 10,000 members, and the allocation it opens on shows app users churning far less than everyone else. Engagement is recorded as two categories, high and low, so there is no binning choice to argue about. Move members between the cells and watch two numbers at once. The pooled gap swings and can change sign. The adjusted difference within engagement is only ever what the third slider says, and it disappears once a group holds no comparison members. Neither number is a measured effect of the app.

Pooled gap (app minus no app)
−6.8 pp
Adjusted difference within engagement
−1.0 pp
02,0004,0006,000members3.0%3,000App4.0%1,000No app13.0%1,000App14.0%5,000No appHigh engagementLow engagement

Bar height is the number of members, on a fixed axis. The bold figure above each bar is that cell's annual churn rate, and an empty cell shows an em dash because nothing is observed in it.

Pooled annual churn, all 10,000 members, bar track fixed from 0% to 15%App users4,000 members5.5%No app6,000 members12.3%
Pooled app churn 5.5%Pooled no-app churn 12.3%Expected app churn (members a year) 220
High-engagement members with the app3,000 of 4,000 (75.0%)
Low-engagement members with the app1,000 of 6,000 (16.7%)
Assumed within-group difference1.0 pp lower
Why engagement is the confounder
EngagementAppChurnunknownEngagement drives both uptake and churnDashed means nothing here measures it

High-engagement members hold the app more often than low-engagement members, 75.0% against 16.7%. Engagement feeds uptake and it feeds churn, which is what makes it a confounder rather than a detail.

Observed difference after holding engagement fixed: −1.0 pp. That figure sits outside the diagram on purpose. It is an adjusted association, so it belongs to the dashed link as a question and not as an answer.

Annual churn by engagement group. Gap is app churn minus no-app churn in percentage points, so a negative gap means app users churn less. Every figure here is an observed difference, never a measured effect, and a group holding members in only one cell shows no gap at all.
GroupApp churnNo-app churnGap (pp)Group size
High engagement
3.0%
n = 3,000
4.0%
n = 1,000
−1.0 pp4,000
Low engagement
13.0%
n = 1,000
14.0%
n = 5,000
−1.0 pp6,000
Pooled
5.5%
n = 4,000
12.3%
n = 6,000
−6.8 pp10,000
Pooled, app users churn at 5.5% against 12.3% for everyone else, a gap of −6.8 pp. Holding engagement fixed, the observed difference in each group is −1.0 pp. The pooled figure splits exactly into that adjusted difference of −1.0 pp plus a composition term of −5.8 pp. The app group is loaded with high-engagement members, who churn at 4.0% before the app is considered at all, so the pooled comparison is partly reporting who holds the app rather than anything the app did. The −1.0 pp that survives the split is an association within engagement, not a measured causal effect. That is why the app to churn link in the diagram is drawn dashed. A randomised holdout or another defensible causal design is needed before anyone forecasts what a campaign delivers.
Stylised model. The churn rate of each cell is held fixed while members are moved between cells, which is an assumption of the teaching case rather than a fitted result. Churn counts are rate-implied expectations rather than a tally of people, so they can land on a fraction. The default allocation reproduces the worked figures on this page.
Correlation and causationOpen in Dr Phil's Quant Lab ↗