Goodhart's Law: What Happens the Moment You Attach Money to a Number
In 1902, colonial administrators in Hanoi decided the city had too many rats. They offered a bounty for each rat killed, payable on presentation of a rat's tail. Tails poured in. The bounty was paid. And the rat population did not fall.
Investigators eventually found rats running around the city with no tails. Enterprising residents had worked out that a live rat with its tail removed would go on to breed more rats — and therefore more tails. Others had begun farming rats outright. The programme was not merely ineffective; it had created an industry devoted to producing the exact thing it was meant to eliminate.
Nobody in this story was stupid. The administrators wanted fewer rats and chose a measurable proxy. The residents responded rationally to the incentive placed in front of them. The failure was structural, and it has a name: when a measure becomes a target, it ceases to be a good measure.
The mechanism: proxies and the gap they hide
Almost nothing we actually care about can be measured directly. "A healthy company," "good teaching," "public safety," "a well-run hospital" — these are real but not countable. So we choose proxies: revenue, test scores, reported crime, mortality rates. Each proxy is chosen precisely because it correlates with the thing we cannot see.
The correlation is genuine at the moment of selection. That is what makes the choice reasonable. But the correlation exists under a specific condition: that people are pursuing the underlying goal, and the metric is passively observing them. Rat tails tracked rat deaths only while nobody was trying to produce tails.
Attach consequences — money, promotion, ranking, funding — and the condition dissolves. Now there are two routes to a higher number: improve the underlying reality, or move the metric directly. The second is almost always faster, cheaper, and lower-risk. Effort migrates accordingly, and the correlation that justified the metric quietly dies.
The critical and frequently missed point is that this does not require dishonesty. Most Goodhart failures involve no rule-breaking at all. A teacher who narrows instruction to tested material is doing their job as defined. A support agent who closes tickets quickly is following policy. The system asked for a number and got one.
The taxonomy of failure
Goodhart failures come in recognisable varieties, and distinguishing them tells you which fix applies.
Selection effects. The metric improves because the population changes, not because performance did. Surgeons ranked publicly on patient survival have a well-documented and entirely rational response: decline the highest-risk cases. Survival rates rise across every published table while the sickest patients — the ones with the most to gain from a skilled surgeon — find no one willing to operate. The measure improved and the goal moved backwards.
Narrowing. The metric captures one dimension of a multi-dimensional goal, so everything unmeasured decays. Teaching to a test is the standard example, but the pattern is universal: measure calls handled per hour and you get short calls, not solved problems. Measure lines of code and you get verbose code. The unmeasured dimensions do not merely stagnate — they are actively sacrificed, because time spent on them is time not spent on the scored dimension.
Direct manipulation. The number is altered without touching reality at all. Reclassifying incidents so they fall outside the reporting category; timing transactions to land in a favourable period; the rat farms of Hanoi.
Goal displacement. The most insidious form, in which the organisation genuinely forgets what the metric was for. After enough years, hitting the number becomes the purpose, and someone who improves the real outcome while missing the target is treated as having failed. At this stage the metric has fully replaced the mission, and people defending the mission look like they are making excuses.
What actually helps
There is no clean escape. You cannot manage what you refuse to measure, and every measure is corruptible. But the failure is manageable, and a few approaches genuinely reduce the damage.
Measure many things, especially the ones that move in opposition. A single metric is trivially gamed; a basket that includes the natural side-effects of gaming is much harder. Pair speed with quality, growth with retention, volume with error rate. Anyone optimising one at the expense of the other becomes visible immediately.
Rotate metrics before they ossify. Gaming strategies take time to develop and institutionalise. A metric that changes periodically never accumulates a mature industry of workarounds. This has real costs in comparability, which is precisely why it is rarely done — and why gaming becomes so sophisticated in metrics that have stood unchanged for decades.
Keep judgement in the loop, and protect it. The strongest defence against Goodhart is a human who understands the actual goal and has authority to override the number. This is unfashionable, because it is subjective and does not scale cleanly. But metrics exist to inform judgement, and systems that eliminate judgement entirely in favour of automated scoring are the most reliably gamed of all.
Watch the gap between the metric and the story. When the numbers improve and the people closest to the work say things are getting worse, believe the people. That divergence is the earliest and most reliable signal that a proxy has detached from its goal.
This connects directly to the principal-agent problem: metrics exist largely because principals cannot observe what agents actually do, and Goodhart's Law describes what happens to any observation the agent can influence. It also explains why the innovator's dilemma is so hard to escape — the metrics that govern an incumbent's decisions are exactly the ones a disruptor is not competing on.
The Hanoi administrators were not fools. They faced an unmeasurable goal, chose a countable proxy, and attached money to it — which is precisely what every organisation does, every quarter, in every performance review. The rats with no tails are still running.