Skip to content

An independent study reference written by Dr Phuc V. Nguyen. It is not official subject material — for assessment requirements always follow your subject outline and vUWS.

Designing a KPI that works

A key performance indicator is a number chosen to steer behaviour, which is what separates it from an ordinary statistic. A workable KPI has an unambiguous definition covering numerator, denominator and time window, an owner who can move it, and a link to a decision someone will actually make. It also needs a guardrail, a second measure that catches the damage the first one invites. The failure mode is predictable. Once a measure carries consequences, people optimise the measure, and any gap between the measure and the goal becomes the thing the organisation pays for.

Try it yourself

The definition is the number

One month of a support desk, one stylised log of 43 events, and a headline that wording alone can move a long way. Set the numerator, the denominator and the window, and every figure is recounted from the log. Nothing about the service changes while you do it.

Reported by this definition
81.8%
Spread across every setting
43.8 pp
rateselected 81.8%0%25%50%75%100%band lowest 56.3%band highest 100.0%

The shaded band is the range every one of the 24 switch settings can produce from this same log at a 28-day window. The thick line is the setting selected on the right.

Over unique customers 81.8%Over sessions 78.6%Over tickets 75.0%

The highest of the three reads 81.8% over unique customers, and the lowest reads 75.0% over tickets opened. A unit is counted as served once any one of its tickets qualifies, so which grain reads highest turns on whether the customers and sessions that came back more than once are the ones being served.

Numerator

This switch is not read while the numerator counts resolutions, because an acknowledgement resolves nothing.

Denominator
Work that finishes after the window closes
Window length28 days (to hour 672)
The log, 43 rows. Hours are offsets from a fixed origin, not a clock.
HourEventTicketSessionCustomerUnder this definition
3ticket openedT01S1C1in window
5first response (agent)T01S1C1not used
14ticket openedT02S2C2in window
15first response (automated)T02S2C2not used
16ticket openedT03S2C2in window
20resolvedT01S1C1counts
28first response (agent)T03S2C2not used
31ticket openedT04S3C3in window
33first response (agent)T04S3C3not used
46resolvedT02S2C2counts
52resolvedT03S2C2counts
58ticket openedT05S4C4in window
59first response (automated)T05S4C4not used
95ticket openedT06S5C5in window
101first response (agent)T06S5C5not used
120ticket openedT07S6C6in window
121first response (automated)T07S6C6not used
126ticket openedT08S6C6in window
130resolvedT06S5C5counts
162ticket openedT09S7C7in window
166first response (agent)T09S7C7not used
172resolvedT09S7C7counts
190resolvedT07S6C6counts
205ticket openedT10S8C8in window
208first response (agent)T10S8C8not used
236resolvedT10S8C8counts
268ticket openedT11S9C1in window
270first response (automated)T11S9C1not used
300resolvedT11S9C1counts
331ticket openedT12S10C4in window
340first response (agent)T12S10C4not used
402ticket openedT13S11C9in window
404first response (agent)T13S11C9not used
430resolvedT13S11C9counts
511ticket openedT14S12C10in window
512first response (automated)T14S12C10not used
553resolvedT14S12C10counts
640ticket openedT15S13C10in window
645first response (agent)T15S13C10not used
665ticket openedT16S14C11in window
690first response (agent)T16S14C11not used
700resolvedT15S13C10counts
712resolvedT16S14C11counts
The desk reports 81.8%, which is 9 of 11 unique customers. Counting the same qualifying work over the three grains gives 81.8% over unique customers, 78.6% over sessions and 75.0% over tickets. The highest of the three reads 81.8% over unique customers, and the lowest reads 75.0% over tickets opened. A unit is counted as served once any one of its tickets qualifies, so which grain reads highest turns on whether the customers and sessions that came back more than once are the ones being served. The automated acknowledgement switch is not read here, because the numerator counts resolutions. Crediting work that finished after the window closed gives 81.8% against 72.7% for counting only what happened inside it. The 2 tickets that opened inside the window and finished after it are the whole argument. Across the 24 settings these switches can take, the same desk over this window reports anywhere from 56.3% to 100.0%, a spread of 43.8 points. The highest is tickets with any first response, automated ones included over unique customers, crediting work that finished after the window closed. The lowest is tickets with a human first response over tickets opened, crediting only what happened inside the window. Neither is dishonest and the desk did the same work in both. That spread is the argument an organisation has when the definition is not written down.
Stylised log. The 43 rows are written for teaching rather than drawn from a real service desk, and the hours are offsets from a fixed origin rather than dates. Every figure above is recounted from those rows by one classifier, so the table markers and the headline cannot disagree.

Why it matters

A call centre told to cut average call time will cut it. Some of that comes from working smarter and some from hanging up on people, and the number itself cannot tell the difference. That is why a measure carrying a bonus needs a companion measure that would move the wrong way if people gamed it. Choose the pair together, never one on its own.

Before you read on — recall

A support team is measured on average time to first response. Which companion measure best protects against the obvious gaming?

Formulas

A metric tree
R=V×c×AˉR = V \times c \times \bar{A}
Revenue RR decomposes into visitors VV, conversion rate cc and average order value Aˉ\bar{A}. With 200,000 visitors, a conversion rate of two and a half per cent and an average order of A$80, revenue is A$400,000. Lifting conversion alone to two point eight per cent takes it to A$448,000. The decomposition matters because each factor has a different owner and a different lever, so it tells you who can act.
Why the denominator is a decision
c=orders in the windowunique visitors in the windowc = \frac{\text{orders in the window}}{\text{unique visitors in the window}}
Swap unique visitors for sessions and the same business reports a different conversion rate, usually a lower one, because a single person may visit several times. Neither definition is wrong. What is wrong is leaving the choice unwritten, because two teams then report different numbers for the same month and the meeting turns into an argument about arithmetic.

Worked examples

Scenario

A logistics company makes on-time delivery percentage the headline KPI for depot managers, with a quarterly bonus attached.

Solution

Within a quarter the number improves. Some of the gain is real. Some comes from quoting longer delivery windows, which makes on-time easier to hit while customers wait longer, and some from recording a delivery as complete when the van reaches the suburb rather than the door. Neither is fraud, and both are rational responses to the measure as written. The repair is to define the promise date as the one quoted to the customer at order time, then pair the KPI with average quoted lead time and with delivery complaints per thousand parcels.

Scenario

A university service desk reports 40 numbers on a monthly dashboard. Leadership says the dashboard is not useful.

Solution

Forty numbers is a reporting habit rather than a set of indicators. Ask which decisions the leadership group actually makes each month, then keep only measures that could change one of them. Usually that leaves three or four, each with a named owner, a written definition and a threshold that triggers a conversation. Everything else moves to a reference page for people who need the detail. A dashboard where everything is important tells a reader nothing about where to look first.

Common mistakes

  • ✗A KPI is simply an important number. It is a number with consequences attached, and consequences change behaviour. Choosing one without asking how it could be met dishonestly leaves the gaming to chance.
  • ✗More indicators give a fuller picture. Past a handful, attention splinters and nothing gets acted on. A small set with clear owners beats a long list nobody can prioritise.
  • ✗A KPI can only be gamed by dishonest people. Most gaming is ordinary and rational. Staff meet the measure they are judged on, and the shortfall appears in whatever the measure failed to capture.
  • ✗Once defined, a KPI should never change, for the sake of comparability. Definitions do need to be stable enough to compare periods, but a measure tied to a strategy that has moved on steers the organisation towards last year goal. Review it deliberately and record the change.

Revision bullets

  • •A KPI carries consequences, which is why it changes behaviour
  • •Write down numerator, denominator, time window and filters
  • •Every KPI needs an owner who can genuinely move it
  • •Pair each KPI with a guardrail that catches the obvious gaming
  • •Decompose a headline metric into factors with different owners
  • •Keep the set small enough that people know where to look first

Quick check

A support team is measured on average time to first response. Which companion measure best protects against the obvious gaming?

Two teams report conversion rates of 2.5 and 3.4 per cent for the same month and the same site. What should be inspected first?

Connected topics

More in Decisions and Models

Sources

  1. Ridgway, V. F. "Dysfunctional Consequences of Performance Measurements." Administrative Science Quarterly, 1(2), 1956.
    Early account of how single, composite and multiple performance measures each distort behaviour once people are judged by them.
  2. Campbell (1979)
    Campbell, D. T. "Assessing the impact of planned social change." Evaluation and Program Planning, 2(1), 1979.
    Argues that the more a quantitative indicator is used for decision making, the more it will be corrupted and the more it will distort what it was meant to monitor.
  3. Strathern (1997)
    Strathern, M. "Improving ratings: audit in the British University system." European Review, 5(3), 1997.
    Source of the widely quoted formulation that a measure ceases to be a good measure once it becomes a target.
  4. Kaplan & Norton (1992)
    Kaplan, R. S., & Norton, D. P. "The Balanced Scorecard: Measures That Drive Performance." Harvard Business Review, 70(1), 1992.
    Argues that a single financial indicator is too narrow to steer an organisation, and proposes a small balanced set of measures instead.
How to cite this page
Dr. Phil's Quant Lab. (2026). Designing a KPI that works. Business Analytics Atlas. https://phucnguyenvan.com/analytics_atlas/concept/ba-kpi-design
Next concept
Experiments and A/B testing
Built by Dr. Phuc V. Nguyen ·Follow on LinkedInWork with PhilEmail