Skip to content
The Operon Library

Volume X · Chapter 10

Healthy Metrics, Dangerous Metrics

Goodhart's law; vendor-metric skepticism; metrics that improve behavior vs. metrics that deform it.2026-07-13 · 8 min read

Nine chapters into this volume, an engineering leader would have a respectable dashboard: decision velocity trending down (good — decisions resolve faster), an AI-suggestion acceptance rate trending up, cost per merged change flattening, a handful of quality and cycle-time signals from the chapters in between. Every line on the chart moves the right direction. The dashboard is, by construction, good news. The uncomfortable question this chapter exists to ask is the one almost nobody asks of their own dashboard: what would it look like if the numbers were improving and the work was not?

It would look exactly like this. A metric that goes up because the underlying reality improved and a metric that goes up because people learned what the number rewards produce an identical chart. The chart cannot tell you which one happened — only a second, independent signal can. This is not a hypothetical hedge to close out a measurement volume with false modesty. It is the single most common failure mode in the history of measuring engineering work, it has a name, and that name is usually misquoted.

A team that got better at the number

The clearest documented case is not from software’s earliest decades — it is a first-hand account from agile coach Joshua Kerievsky, published on the Industrial Logic blog in October 2012. Kerievsky’s firm had spent four and a half years coaching a large organization near Tampa, Florida. One twenty-five-person team had been performing well under ordinary agile practice, averaging around fifty-two story points per iteration. Then a director leaned on the team to “go faster.” Within weeks, the same team’s velocity climbed into the high eighties — a jump with no corresponding change in scope, staffing, or delivered functionality. When Kerievsky asked a team member what had changed, the answer was blunt: “These days around here if you sneeze, you get a story point.”

Nothing about the team’s output changed. What changed was that velocity had quietly become a target instead of a measurement, and the team — under pressure, not out of malice — found the shortest path to satisfying it. Kerievsky cites this specific episode as the moment he stopped trusting story points as a management signal. It is a small, local story, not a controlled study, and it should be read with that scope in mind. But it is a named person, a named engagement, a specific quote, not a folklore anecdote passed down without a source — which is precisely the standard this chapter is holding every other metric in this volume to, including its own.

Goodhart’s law, precisely

The mechanism the Tampa team lived through has a name, and the name is attributed slightly wrong more often than it is attributed right. Economist Charles Goodhart stated the original idea in 1975, writing about UK monetary policy: “Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.” Goodhart was making a narrow, technical point about central banks — once a bank starts targeting a specific measure of the money supply to control inflation, that measure stops behaving the way it did when it was merely observed, because the relationships that made it a useful indicator assumed nobody was steering by it.

The crisp, universally quoted version — “when a measure becomes a target, it ceases to be a good measure” — is not Goodhart’s own wording. It is anthropologist Marilyn Strathern’s 1997 paraphrase, published in a paper on the British university audit system, generalizing an idea that management theorist Keith Hoskin had already started extending beyond economics the year before. Strathern took Goodhart’s narrow claim about monetary aggregates and reframed it as a claim about any human system under evaluation — audits, rankings, targets, and, by extension, engineering dashboards. Both attributions matter and they are not interchangeable: Goodhart’s law is the economist’s name for the phenomenon; the sentence engineers actually quote is Strathern’s.

Goodhart described why a targeted statistic breaks. Strathern gave the break a sentence anyone could repeat in a standup.

Software engineering has its own founding version of this story, already familiar from this volume’s opening chapter: the lines-of-code era. IBM’s System/360 project in the 1960s counted lines of code because a punch card was a natural, countable unit of visible programmer output — Fred Brooks documented the resulting productivity accounting, and its perversities, in The Mythical Man-Month. Once lines of code became something a programmer was implicitly rewarded for producing, verbosity stopped being a cost and became a strategy. The metric did not lie about what it measured. It measured typing. It was never a good proxy for the thing anyone actually wanted, which was working software, and pressure on the number made the gap worse, not better.

Why vendor metrics deserve extra scrutiny

AI coding tool vendors did not invent this problem, but they inherit its sharpest edge. A vendor’s business depends on demonstrating that its product drives adoption and output, and the metrics cheapest for a vendor to instrument and publish — suggestions shown, suggestions accepted, lines generated — sit at the base of the same pyramid this volume has returned to repeatedly: the layer that is easiest to observe and least connected to whether the work was any good. This is a structural incentive, not an accusation of dishonesty against any specific company. A vendor optimizing its own published number is doing exactly what Goodhart described, because engagement metrics are the number closest to its hand and the number investors and buyers ask about first.

The gap this produces is visible in independent data, not just in argument. Stack Overflow’s 2025 Developer Survey found more developers actively distrust the accuracy of AI tool output (46%) than trust it (33%) — even as usage kept climbing toward 84% adoption — and that positive sentiment toward AI tools fell from over 70% in 2023 and 2024 to 60% in 2025. GitClear’s analysis of more than 211 million changed lines found code duplication rising sharply and the refactoring share of changes collapsing over the same period AI-generated code became common, though the methodology is contested and should be read as a documented tension rather than a settled fact. Neither of those signals came from a vendor dashboard. That is the pattern worth internalizing: the research a reader should weight most heavily is the research an interested party did not produce and could not have shaped — the METR randomized trial this volume has already leaned on, an independent survey of tens of thousands of developers, a peer-scrutinized dataset — not the number that arrives pre-formatted in a vendor’s quarterly product update.

None of this means vendor telemetry is worthless. It means a number a vendor is incentivized to see rise deserves the same skepticism a number your own team is incentivized to see rise deserves — and no less.

A checklist, not a verdict

No metric is immune to Goodhart’s law — that is the wrong bar. The useful question for any metric on a dashboard is narrower: does gaming it cost roughly as much as doing the real work, or is gaming it cheaper? A metric only stays honest as long as satisfying it and actually improving the underlying outcome remain the same act. The following questions are not a scoring rubric with a pass threshold; they are the questions worth asking out loud, in the room, before a metric goes on a dashboard someone will be evaluated against.

PropertyHealthy signalDangerous signal
Cost of gamingSatisfying the number requires doing most of the real work anywaySatisfying the number is cheaper than doing the real work
IsolationReported alongside at least one independent, harder-to-fake signalReported alone, on its own chart, with nothing to contradict it
Direction reflexA move in the number without a matching qualitative check is treated as inconclusiveA rising number is treated as success by itself
Who is being measuredThe people closest to the work helped choose the metricThe metric was imposed on the people who now optimize for it
LifespanRetired or reweighted once it stops correlating with the outcome it proxiesKept indefinitely because it is already wired into a dashboard or a bonus

Turning the checklist on this volume

A chapter that applies this framework only to lines of code and other people’s vendors is not a serious application of it. Two metrics this volume has already proposed fail parts of its own checklist, and saying so plainly is more useful than pretending they don’t.

Decision velocity, introduced in Chapter 4 as the time between a decision surfacing and its being recorded, has a cheap failure mode: a team under pressure to show a faster number can log vague, premature decisions to keep the clock moving, the same way the Tampa team logged inflated story points. A decision recorded quickly and a decision made well are not the same event, and decision velocity on its own cannot tell them apart. The fix the checklist itself recommends is pairing the number with an independent, harder-to-fake companion — a decision reversal rate, or a periodic sample of logged decisions checked against what actually shipped — so a velocity improvement that comes from rushing shows up as a reversal-rate cost somewhere else on the same dashboard.

Acceptance rate, introduced in Chapter 6 as a proxy for whether AI suggestions were useful, sits even lower on the pyramid than decision velocity — it is closer to lines-of-code territory than this volume would like. Accepting a suggestion is an activity metric: it records that a keystroke happened, not that the suggestion survived review, shipped, or worked. The same critique this chapter leveled at vendor-published acceptance numbers applies with equal force to an internal team tracking its own. The honest pairing is retention after some fixed horizon — whether the accepted code was still there, largely unmodified, N days later — or a defect rate attributed specifically to AI-authored lines. Acceptance rate answers a narrower question than it is usually asked to answer, and this volume should not exempt its own chapter from that observation just because it wrote the metric down first.

What survives contact with the number

The honest conclusion is not that some cleverer metric exists which resists gaming entirely. None does, including the ones proposed elsewhere in this volume. What survives is not a metric but a practice: several metrics triangulated against each other rather than reported alone, at least one of them paired with a periodic qualitative review a spreadsheet cannot fake, and an explicit willingness to retire a number once it stops tracking the outcome it was built to proxy. A team that treats its dashboard as permanent has already lost the argument with Goodhart’s law; a team that treats every metric on it as provisional, subject to review, and replaceable has at least given itself a chance. Building that practice — not selecting one more clever number — is the subject the volume closes with next.

For Discussion

  1. Pick the metric your team is most proud of improving this quarter. If someone on the team wanted to move that number without improving the underlying outcome, could they — and would you notice within a month if they had?
  2. Which of your current engineering metrics came from a vendor’s own dashboard rather than from data your team controls end to end? What would change about how much you trust it if it came from an independent source instead?
  3. Name one metric your team has never retired, even though nobody could tell you the last time it changed a decision. What is it still doing on the dashboard?

References

  1. establishedOriginal 1975 formulation: "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes"Charles Goodhart, "Problems of Monetary Management: The U.K. Experience" · 1975
  2. establishedThe popularized paraphrase — "When a measure becomes a target, it ceases to be a good measure" — is Strathern’s generalization of Goodhart, not his own wordingMarilyn Strathern, "‘Improving ratings’: audit in the British University system," European Review, vol. 5 · 1997
  3. emergingDocumented first-hand case: a 25-person team’s story-point velocity jumped from ~52 to the high 80s within weeks of being told to "go faster," with no change in delivered scopeJoshua Kerievsky, Industrial Logic blog, "Stop Using Story Points" · 2012-10-12
  4. establishedFred Brooks documents early lines-of-code productivity accounting and its perversities on the OS/360 projectFrederick P. Brooks Jr., The Mythical Man-Month (Addison-Wesley) · 1975
  5. establishedMore developers actively distrust AI tool output accuracy (46%) than trust it (33%); positive sentiment fell from 70%+ in 2023–2024 to 60% in 2025, despite 84% adoptionStack Overflow Developer Survey 2025 · 2025-07
  6. contestedRising code duplication and collapsing refactoring share across 211M+ changed lines in the AI-adoption period (methodology contested)GitClear research · 2025-01
  7. establishedAI functions as an amplifier of existing organizational strengths and dysfunctions rather than a uniform productivity multiplierDORA — State of AI-assisted Software Development 2025 · 2025-09
  8. establishedRandomized controlled trial: experienced developers measured 19% slower with early-2025 AI tools while estimating themselves ~20% faster — the reference case for trusting independently replicated research over self-reported or vendor-reported productivity claimsMETR · 2025-07-10