A screening model scores 1,000 applicants in each of two groups. 300 people in group A would succeed in the role and 250 in group B, so the base rates differ, 30.0% against 25.0%. Move a threshold and the four measures below are recomputed from the counts above it. This widget never returns a single fair or unfair verdict, because the criteria measure different things and choosing between them is a business decision about which harm to accept.
The shaded strip around zero is an illustrative ±1.0 pp band drawn for this widget only. No standard sets it, and nothing here treats a bar inside it as a pass.
Both panels use the same fixed count axis, up to 240 applicants in a band, so the two shapes stay comparable while the thresholds move. Bands at or above the dashed marker are shortlisted and are drawn at full strength, the rest are faded.
equal precision. 0.0 pp, inside the illustrative ±1.0 pp band.
Equalised odds
equal true positive rate and equal false positive rate. +24.0 pp and +1.5 pp. Both have to close together, so this criterion reads the two error-rate bars as a pair rather than one of them.
Base rate A 30.0%Base rate B 25.0%Base-rate gap +5.0 pp
Rates at the current thresholds, each printed with the counts it comes from. Gaps are computed from those counts and not from the rounded percentages, so a gap can differ in the last printed digit from subtracting the two rounded rates. A rate with a zero denominator is undefined and shows an em dash.
Measure
Group A (band 8 and above)
Group B (band 8 and above)
Gap (A − B)
Selection rate
shortlisted / applicants
20.0%
200 / 1,000
10.0%
100 / 1,000
+10.0 pp
True positive rate
shortlisted who would succeed / everyone who would succeed
60.0%
180 / 300
36.0%
90 / 250
+24.0 pp
False positive rate
shortlisted who would not succeed / everyone who would not
2.9%
20 / 700
1.3%
10 / 750
+1.5 pp
Precision
shortlisted who would succeed / shortlisted
90.0%
180 / 200
90.0%
90 / 100
0.0 pp
Group A, threshold band 8 and above. Counts of applicants.
Decision
Would succeed
Would not
Total
Shortlisted
180
20
200
Not shortlisted
120
680
800
Total
300
700
1,000
Group B, threshold band 8 and above. Counts of applicants.
Decision
Would succeed
Would not
Total
Shortlisted
90
10
100
Not shortlisted
160
740
900
Total
250
750
1,000
Group A shortlists 200 of 1,000 applicants and 180 of them would succeed, precision 90.0%, true positive rate 60.0%. Group B shortlists 100 of 1,000 applicants and 90 of them would succeed, precision 90.0%, true positive rate 36.0%. Exactly zero at this setting: the precision gap. Precision matches exactly, 90.0% in both groups, and the true positive rate does not. 60.0% of the 300 people in group A who would succeed are shortlisted, against 36.0% of the 250 in group B. A team that checks precision alone passes this model, and Group B is the group it reaches least. The result to carry away is the one on this page. When group base rates differ, equal precision and equal error rates cannot both hold for a classifier that still makes mistakes and still separates people, provided the shared true positive rate and the shared false positive rate are both above zero. That proviso matters, and there are two boundaries rather than one. A rule that never produces a false positive has a precision of one in every group whatever the base rate, and buys that agreement by missing almost everyone it should find. A rule that never produces a true positive has a precision of zero in every group, and flags only people it is wrong about. On the two populations here both boundaries are out of reach with anyone shortlisted, because every score band of both groups holds people who would succeed and people who would not. Away from those boundaries the conflict is real, and Chouldechova and, independently, Kleinberg, Mullainathan and Raghavan proved it. Choosing among the criteria is a decision about which harm to accept, and it belongs to the business, not to the modeller alone.
No threshold on this page repairs a biased label, a biased sample or a historical outcome already recorded in the training data. Those sit upstream of every number here. The label can be the whole problem: a widely deployed United States health algorithm ranked patients by predicted spending rather than illness, which reads a group that historically received less care as healthier. Obermeyer and colleagues documented it in 2019. Nothing was wrong with the code. The bug was the choice of target variable.
Stylised model. The two score distributions are invented for teaching and built so that one shared threshold at band 8 reproduces the shortlist figures on this page exactly, 200 shortlisted in group A of whom 180 would succeed and 100 in group B of whom 90 would. They are not fitted to any data.