Skip to content

An independent study reference written by Dr Phuc V. Nguyen. It is not official subject material — for assessment requirements always follow your subject outline and vUWS.

Algorithmic bias

A model is biased when the errors it makes, the outcomes it allocates or the burdens it imposes differ systematically across groups of people. The cause is rarely a column labelled with a protected attribute. Bias enters through the label you chose to predict, which is usually a convenient proxy for the thing you actually care about; through the sample, meaning who is in the training data and who is missing; and through historical outcomes, because the data records what the organisation used to do and a model that fits it well reproduces it faithfully. Fairness has several formal definitions that contradict one another, so "make it fair" is not a single instruction.

Try it yourself

Fairness criteria measured separately

A screening model scores 1,000 applicants in each of two groups. 300 people in group A would succeed in the role and 250 in group B, so the base rates differ, 30.0% against 25.0%. Move a threshold and the four measures below are recomputed from the counts above it. This widget never returns a single fair or unfair verdict, because the criteria measure different things and choosing between them is a business decision about which harm to accept.

Gap = group A minus group B, percentage points, fixed scale ±30 pp-30-150+15+30Selection rate+10.0 ppTrue positive rate+24.0 ppFalse positive rate+1.5 ppPrecision0.0 pp

The shaded strip around zero is an illustrative ±1.0 pp band drawn for this widget only. No standard sets it, and nothing here treats a bar inside it as a pass.

Group A200 shortlisted, 180 would succeedshortlisted →12345678910applicantsGroup B100 shortlisted, 90 would succeedshortlisted →12345678910model score band, 10 is the strongestapplicantswould succeedwould not succeeddashed line = threshold

Both panels use the same fixed count axis, up to 240 applicants in a band, so the two shapes stay comparable while the thresholds move. Bands at or above the dashed marker are shortlisted and are drawn at full strength, the rest are faded.

Shared threshold, both groupsband 8 and above
The four criteria, reported separately
Demographic parity
equal selection rate. +10.0 pp, outside the illustrative ±1.0 pp band.
Equal opportunity
equal true positive rate. +24.0 pp, outside the illustrative ±1.0 pp band.
Predictive parity
equal precision. 0.0 pp, inside the illustrative ±1.0 pp band.
Equalised odds
equal true positive rate and equal false positive rate. +24.0 pp and +1.5 pp. Both have to close together, so this criterion reads the two error-rate bars as a pair rather than one of them.
Base rate A 30.0%Base rate B 25.0%Base-rate gap +5.0 pp
Rates at the current thresholds, each printed with the counts it comes from. Gaps are computed from those counts and not from the rounded percentages, so a gap can differ in the last printed digit from subtracting the two rounded rates. A rate with a zero denominator is undefined and shows an em dash.
MeasureGroup A (band 8 and above)Group B (band 8 and above)Gap (A − B)
Selection rate
shortlisted / applicants
20.0%
200 / 1,000
10.0%
100 / 1,000
+10.0 pp
True positive rate
shortlisted who would succeed / everyone who would succeed
60.0%
180 / 300
36.0%
90 / 250
+24.0 pp
False positive rate
shortlisted who would not succeed / everyone who would not
2.9%
20 / 700
1.3%
10 / 750
+1.5 pp
Precision
shortlisted who would succeed / shortlisted
90.0%
180 / 200
90.0%
90 / 100
0.0 pp
Group A, threshold band 8 and above. Counts of applicants.
DecisionWould succeedWould notTotal
Shortlisted18020200
Not shortlisted120680800
Total3007001,000
Group B, threshold band 8 and above. Counts of applicants.
DecisionWould succeedWould notTotal
Shortlisted9010100
Not shortlisted160740900
Total2507501,000
Group A shortlists 200 of 1,000 applicants and 180 of them would succeed, precision 90.0%, true positive rate 60.0%. Group B shortlists 100 of 1,000 applicants and 90 of them would succeed, precision 90.0%, true positive rate 36.0%. Exactly zero at this setting: the precision gap. Precision matches exactly, 90.0% in both groups, and the true positive rate does not. 60.0% of the 300 people in group A who would succeed are shortlisted, against 36.0% of the 250 in group B. A team that checks precision alone passes this model, and Group B is the group it reaches least. The result to carry away is the one on this page. When group base rates differ, equal precision and equal error rates cannot both hold for a classifier that still makes mistakes and still separates people, provided the shared true positive rate and the shared false positive rate are both above zero. That proviso matters, and there are two boundaries rather than one. A rule that never produces a false positive has a precision of one in every group whatever the base rate, and buys that agreement by missing almost everyone it should find. A rule that never produces a true positive has a precision of zero in every group, and flags only people it is wrong about. On the two populations here both boundaries are out of reach with anyone shortlisted, because every score band of both groups holds people who would succeed and people who would not. Away from those boundaries the conflict is real, and Chouldechova and, independently, Kleinberg, Mullainathan and Raghavan proved it. Choosing among the criteria is a decision about which harm to accept, and it belongs to the business, not to the modeller alone.
No threshold on this page repairs a biased label, a biased sample or a historical outcome already recorded in the training data. Those sit upstream of every number here. The label can be the whole problem: a widely deployed United States health algorithm ranked patients by predicted spending rather than illness, which reads a group that historically received less care as healthier. Obermeyer and colleagues documented it in 2019. Nothing was wrong with the code. The bug was the choice of target variable.
Stylised model. The two score distributions are invented for teaching and built so that one shared threshold at band 8 reproduces the shortlist figures on this page exactly, 200 shortlisted in group A of whom 180 would succeed and 100 in group B of whom 90 would. They are not fitted to any data.

Why it matters

A model learns to repeat the pattern in the data it was handed. If that data records what an organisation used to do, the model learns to do the same thing faster, more cheaply and with a much straighter face. It is not prejudiced. It is obedient. That is why deleting the gender column changes almost nothing when twenty other columns carry the same signal between them.

Before you read on — recall

A lender removes gender from its credit model and confirms the coefficient is gone. Approval rates for women remain 9 percentage points below men's. What is the most likely explanation?

Formulas

Demographic parity
P(Y^=1∣A=a)  =  P(Y^=1∣A=b)P(\hat{Y}=1 \mid A=a) \;=\; P(\hat{Y}=1 \mid A=b)
Equal selection rate across groups aa and bb. It ignores whether the groups differ in the outcome being predicted, so enforcing it can mean selecting people the model scores lower.
Equal opportunity (equal true positive rate)
P(Y^=1∣Y=1, A=a)  =  P(Y^=1∣Y=1, A=b)P(\hat{Y}=1 \mid Y=1,\, A=a) \;=\; P(\hat{Y}=1 \mid Y=1,\, A=b)
Among the people who genuinely have the outcome, the same share is identified in each group. This is usually the criterion that matters when the cost of a miss falls on the individual rather than the organisation.
Predictive parity (equal precision)
P(Y=1∣Y^=1, A=a)  =  P(Y=1∣Y^=1, A=b)P(Y=1 \mid \hat{Y}=1,\, A=a) \;=\; P(Y=1 \mid \hat{Y}=1,\, A=b)
A positive flag means the same thing whichever group it lands on. When base rates differ across groups, this cannot hold at the same time as equal false positive and false negative rates, provided the shared true positive rate and the shared false positive rate are both above zero. Those two boundaries are where the conflict disappears, and both are useless as rules. A model that never produces a false positive has a precision of one in every group whatever the base rate, and buys that by missing almost everyone. A model that never produces a true positive has a precision of zero in every group, and flags only people it is wrong about. Away from those boundaries the conflict is real, and it was proved independently by Chouldechova and by Kleinberg, Mullainathan and Raghavan.

Worked examples

Scenario

A screening model shortlists graduate applicants. The team checks it for bias, finds equal precision in both groups, and declares it fair.

Solution

Run the arithmetic. Group A has 1,000 applicants of whom 300 would succeed in the role. The model shortlists 200 and 180 of those would succeed. Group B has 1,000 applicants of whom 250 would succeed. The model shortlists 100 and 90 of those would succeed. Precision is 180/200 and 90/100, so 90 per cent in both groups, and the team's test passes. But the true positive rate is 180/300, or 60 per cent, for A and 90/250, or 36 per cent, for B. The model is equally trustworthy when it says yes, and it never sees most of the capable people in B.

Scenario

A health system ranks patients for a care-management programme using predicted future health spending, because spending data is complete, timely and easy to obtain.

Solution

Spending measures how much care a person received, not how sick they are. Where a group has historically received less care at the same level of illness, the model reads that group as healthier and ranks them lower, which withholds exactly the extra help that would have closed the gap. Obermeyer and colleagues documented this in a widely deployed US algorithm in 2019, and switching the label from cost to a direct measure of illness sharply increased the number of Black patients flagged for additional care. Nothing was wrong with the code. The bug was the choice of target variable.

Common mistakes

  • ✗Removing protected attributes makes a model fair. Postcode, school, name, occupation, purchase history and browsing behaviour all correlate with protected attributes, so the model reconstructs them from what is left. Deleting the column removes your ability to audit, not the model's ability to discriminate.
  • ✗Bias is a data problem, so cleaner data fixes it. Some of it does live in the sample, and coverage gaps are real. The most damaging kind lives in the label: you chose to predict spending, arrests or clicks because they were measurable, and the model is faithful to that choice. No amount of cleaning fixes the wrong target.
  • ✗A fair model is one that satisfies the fairness metric. There are several incompatible definitions, and when group base rates differ, equal precision and equal error rates cannot both hold. Choosing among them is a decision about which harm you are willing to accept, and it belongs to the business, not to the modeller alone.
  • ✗If the model beats the humans it replaces, it is an improvement. A model applies one rule to everyone at once, so a modest systematic error becomes a consistent, scaled and hard-to-contest one. Human decisions are noisier but their inconsistency also leaves more room for a person to be heard.

Revision bullets

  • •Bias enters via the label, the sample and historical outcomes, not a protected column
  • •Dropping a protected attribute leaves the proxies intact and removes the audit
  • •Demographic parity, equal opportunity and predictive parity are different tests
  • •Different base rates plus an imperfect model means those tests cannot all pass
  • •Obermeyer et al. (2019): cost as a proxy for need read sicker patients as healthier
  • •Choosing the fairness criterion is a decision about which harm is acceptable

Quick check

A lender removes gender from its credit model and confirms the coefficient is gone. Approval rates for women remain 9 percentage points below men's. What is the most likely explanation?

A recidivism tool is equally well calibrated for two groups, so a score of 7 implies the same reoffending probability for anyone, and the two groups reoffend at different underlying rates. The tool is accurate but not perfect. Critics show it produces a higher false positive rate for one group. Who is right?

Connected topics

More in Ethics and Governance

Sources

  1. Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. "Dissecting racial bias in an algorithm used to manage the health of populations." Science, 366(6464), 447-453, 2019.
    The clearest published case of label bias: health cost used as a proxy for health need, and what changed when the label was replaced.
  2. Buolamwini, J., & Gebru, T. "Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification." Proceedings of Machine Learning Research, 81, 77-91, 2018.
    Audited three commercial classifiers and found error rates under one per cent for lighter-skinned men and up to roughly a third for darker-skinned women, traced to who appeared in the training and benchmark images.
  3. Chouldechova, A. "Fair Prediction with Disparate Impact: A Study of Bias in Recidivism Prediction Instruments." Big Data, 5(2), 153-163, 2017.
    Shows that when base rates differ, a predictor cannot be both calibrated and equal in false positive and false negative rates.
  4. Kleinberg, J., Mullainathan, S., & Raghavan, M. "Inherent Trade-Offs in the Fair Determination of Risk Scores." Innovations in Theoretical Computer Science (ITCS), 2017.
    The independent formal statement of the same incompatibility, framed as three conditions that cannot hold together except in degenerate cases.
How to cite this page
Dr. Phil's Quant Lab. (2026). Algorithmic bias. Business Analytics Atlas. https://phucnguyenvan.com/analytics_atlas/concept/ba-algorithmic-bias
Next concept
Why ethics belongs in analytics
Built by Dr. Phuc V. Nguyen ·Follow on LinkedInWork with PhilEmail