Skip to content
Accuracy and the class you act on

A month brings 10,000 customer reviews. A model gives each review a score for how negative it looks, and a decision threshold turns that score into an alert for the operations team. Move the threshold and watch two figures that answer different questions: overall accuracy, and how many real complaints reach a human.

This threshold (0.750)
97.5%
96 of 300 complaints caught
Label everything positive
97.0%
0 of 300 complaints caught

This threshold calls 9,752 of the 10,000 reviews correctly against 9,700 for the rule that flags nothing, 52 more, and it reaches 96 of the 300 complaints against none.

0597 reviewsreviews per bandflagged →← not flagged0.000.250.500.751.00model score, 0 reads as positive and 1 reads as negative

negative reviews (complaints), 300, drawn above the line. positive reviews, 9,700, drawn below the line. A faded band sits below the threshold and is not flagged. Both halves share one scale of reviews per pixel. The complaint side peaks at 19 reviews against 597 reviews on the positive side, so it is a thin strip. That is what a rare class looks like, and it is why a measure built on all 10,000 reviews barely notices it.

Precision (negative class) 68.6%Recall (negative class) 32.0%F1 (negative class) 0.436
Precision = 96 caught of 140 flagged. Recall = 96 caught of 300 negative reviews. F1 = 2TP / (2TP + FP + FN) = 192 / 440 = 0.436. That count form gives the same number as 2PR / (P + R) wherever precision is defined.
Decision threshold0.750
Reviews that are genuinely negative3% (300 of 10,000)
Flagged reviews the team can read1,000 a month
The 140 alerts fit inside the stated capacity of 1,000 a month, with room for 860 more.

The lowest threshold that fits 1,000 alerts a month is 0.550. It raises 935 alerts and catches 234 of the 300 complaints, a recall of 78.0%.

Confusion matrix at a threshold of 0.750. Rows are what the model did, columns are what the review actually was, and every cell is a whole number of reviews. The four cells sum to 10,000. Being flagged is not the same as being read. All 140 alerts fit inside what the team can read, so the flagged row is the row that reaches a human and the 9,860 not flagged never do.
ModelActually negativeActually positiveRow total
Flagged as negative
96
complaints caught
44
false alarms
140
alerts
Not flagged
204
complaints missed
9,656
correctly left alone
9,860
not flagged
Column total3009,70010,000
Accuracy is 97.5%, which is 9,752 of 10,000 reviews called correctly. Precision on the negative class is 68.6%, which is 96 of the 140 flagged, and recall is 32.0%, which is 96 of the 300 negative reviews. The alert list holds 96 real complaints mixed with 44 false alarms, and 204 complaints are never flagged at all. On accuracy this threshold calls 52 more reviews correctly than the rule that flags nothing, 9,752 against 9,700 of 10,000, and it surfaces 96 of the 300 complaints against none. Both figures are reported because either one on its own hides the other. The 140 alerts fit inside the stated capacity of 1,000 a month, with room for 860 more. The lowest threshold that fits 1,000 alerts a month is 0.550. It raises 935 alerts and catches 234 of the 300 complaints, a recall of 78.0%. Negative reviews are 3 per cent of the batch, so the 300 negative reviews carry 3 per cent of the accuracy figure and the 9,700 positive ones carry the rest. A rule that never flags anything loses every complaint and still scores 97.0%. Sweeping the threshold, the highest accuracy available here is 97.5% at a threshold of 0.775, and that setting catches 79 of the 300 complaints on 104 alerts. It prints the same figure as the setting on screen at one decimal place, and the exact counts are 9,754 reviews called correctly there against 9,752 here. Compare it with the setting on screen before deciding which measure the choice should turn on.
Stylised scores. The two distributions come from a fixed seeded draw, one shape for negative reviews and one for positive reviews, and no real model, corpus or product is involved. Nothing here claims that any particular method reaches any of these figures. What is real is the arithmetic. The confusion matrix, the rates and the alert volume all follow exactly from the whole-review counts shown. The class shapes never change, so the balance slider moves only how many reviews are allotted to each shape.