Skip to content

Campbell's Law: How High-Stakes Metrics Corrupt the Work

Scribelet Team
12 min read

An engineering org decided that deploy frequency was the number that mattered. It was a reasonable instinct. Teams that ship often tend to ship small, safe, well-tested changes, so the rate of deploys really did track something healthy. Then it became the headline on the leadership dashboard, tied to quarterly targets and praised in all-hands. Within two quarters deploys had tripled, and almost none of them mattered. Engineers split one change into five commits and shipped them separately. Risky, valuable work that needed a careful rollout got deferred, because a big deploy that might fail looked worse on the chart than three trivial ones that could not. Reliability, the thing deploy frequency was supposed to protect, quietly got worse. The dashboard had never looked better.

That is not simply a gamed metric. The number was gamed, yes, but something worse happened underneath it: the actual work changed shape to serve the number, and the work got worse. The team was not measuring shipping anymore. It was measuring its own skill at producing deploys, and that skill had crowded out the shipping. There is a law that names this exact failure, and it is sharper and older than the one most people reach for. It is Campbell's law, and the thing it predicts is not just that your metric will lie to you but that the process the metric was watching will bend itself out of true. Knowing it is the difference between asking "is this number still honest?" and asking the more important question: "what is this number doing to the work?"

What is Campbell's law?

Campbell's law states: the more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures, and the more apt it will be to distort and corrupt the social processes it is intended to monitor. It was formulated in 1976 by the American social scientist Donald T. Campbell, who spent his career on research methodology and watched the same pattern play out across education, policing, and public policy. The classic case is standardized testing: once school funding and teacher jobs ride on test scores, schools start teaching to the test, narrowing the curriculum and, in the worst cases, manipulating the results, so the scores rise while the education they were meant to reflect declines.

Two things in that statement do the real work, and both are easy to skim past. The first is high-stakes. Campbell did not say a measure goes bad the moment you look at it; he said it corrupts in proportion to how heavily it is used for consequential decisions. A number watched out of curiosity stays mostly honest. A number wired to funding, promotions, or public rankings is under enormous pressure, and it decouples fast. The second is the word distort. The measure does not just stop being accurate; the underlying activity it monitors gets deformed to produce the number. The scores go up and the learning goes down. That second clause is why this law deserves its own name, and it is worth keeping the distinction written down somewhere durable, because it is the part everyone forgets first.

Why the stakes set the speed

Nothing about a metric changes when you raise the stakes on it. What changes is how hard people push on it, and Campbell's insight is that the push scales with the consequences. Give a number mild attention and it stays roughly correlated with the thing you care about, because the cheapest way to nudge it is still to do the real work. Attach the number to someone's bonus, their team's budget, or a leaderboard the whole company sees, and the calculus flips. Now there is a strong, direct incentive to move the number by whatever route is cheapest, and the cheapest route is almost never the substance it was a proxy for.

A line chart with the horizontal axis labeled how much the metric drives high-stakes decisions, rising from watched to wired to pay and funding. A reported-metric line climbs steadily across the whole axis. A real-outcome line tracks it closely at low stakes, then crosses a marked corruption threshold where it flattens and falls while the reported line keeps rising, and the widening shaded gap between them is labeled distortion.

The chart above is the shape of the law. On the left, where the metric is only watched, the reported number and the real outcome move together, because moving the number honestly means improving the work. As you move right and the stakes climb, the two lines split at a threshold, and past it they move in opposite directions: the reported metric keeps rising because people are now optimizing it directly, while the real outcome bends downward because the effort and the shortcuts went into the proxy. The shaded gap is the distortion Campbell warned about, and the crucial point is that the gap grows with the stakes, not with time. You can open it in a single quarter by attaching a big enough consequence to a soft enough number. This is also why the better your enforcement, the faster the rot: a metric nobody is rewarded for drifts slowly, and a metric everyone is judged on corrupts almost on contact.

People are not cheating when this happens, or not usually. They are doing the rational thing, which is responding to the incentive actually in front of them, the same way a chain of locally sensible choices can add up to a bad outcome nobody picked. The deploy-frequency team was not sabotaging the company. Each engineer was making the individually correct move given what the org had decided to reward, and the sum of those correct moves was a corrupted process.

Campbell's law vs Goodhart's law

These two get used interchangeably, and they are close cousins, but they are not the same statement and the difference is useful. The short version many people quote, "when a measure becomes a target, it ceases to be a good measure," is Goodhart's law in the phrasing the anthropologist Marilyn Strathern gave it. Campbell's is older in spirit and aimed at a different scale.

Campbell's law (1976)Goodhart's law (1975)
Field of originSocial science, program evaluationEconomics, monetary policy
The core claimA metric used for high-stakes decisions corrupts and distorts the process it monitorsWhen a measure becomes a target, it stops being a good measure
What it emphasizesThe stakes, and damage to the underlying activityThe decoupling of the measure from the goal
Scale it describes bestInstitutions and systems: schools, agencies, whole teamsAny measure under optimization pressure, from one person up
The sharpest one-linerThe higher the stakes, the faster the number lies and the work rotsWhat gets measured as a target gets gamed
Reach for it whenYou need to explain why the work itself got worse, not just the dashboardYou need the clean, general statement of the trap

The practical read is that Goodhart tells you that a targeted metric goes bad, and Campbell tells you why it gets worse the more it matters and that the real damage lands on the activity, not the spreadsheet. There is a third sibling worth knowing, the cobra effect, where an incentive produces the opposite of its intent (a bounty on dead cobras leads to people breeding cobras for the bounty). The cobra effect is the extreme case of both laws: the metric did not just decouple, it inverted. When you are diagnosing a broken incentive, Goodhart names the mechanism, Campbell names the stakes and the collateral damage, and the cobra effect is the punchline when it has gone all the way.

Where Campbell's law shows up in software and knowledge work

Modern teams run almost entirely on proxies, and the high-stakes ones are exactly the ones Campbell's law comes for. Here are the measures that most reliably corrupt the work when the stakes on them rise, what people do instead, and what gets damaged underneath.

High-stakes metricWhat people optimizeThe process it corrupts
Deploy frequency as a headline KPITrivial, split-up deploys; deferred risky workShipping real value, and reliability itself
Story points or velocity tied to reviewsPoint inflation; padded estimatesHonest planning and forecasting
Tickets closed per engineerSplitting work, closing hard tickets as "cannot reproduce"Actually resolving the user's problem
Code review approval rate or speedRubber-stamp approvalsThe review catching real defects
Bug count as a team scorecardNot filing bugs; reclassifying themThe honest record of what is broken
Support time-to-first-responseFast, empty replies inside the SLAThe customer actually getting helped
An AI model's benchmark scoreTuning to the benchmark's quirksThe capability the benchmark stood for

The through-line is Campbell's second clause. In every row it is not only the number that suffers; it is the thing underneath. A team gaming deploy frequency ships worse software. A team gaming ticket counts leaves users unhelped. The dashboard improves and the work rots, which is precisely the trap, because the people above the dashboard see only the line going up. This is the same dynamic by which any behavior you can actually observe becomes the thing people build toward, whether or not it was what you meant to ask for. Raise the stakes on an observable proxy and you do not just get a gamed proxy; you get a workforce quietly optimizing the proxy at the expense of the job.

Which of your metrics are in the danger zone

Campbell's law is not a reason to stop measuring, but it is a reason to look hard at which numbers you have loaded with consequences. A metric is dangerous roughly in proportion to how these line up.

QuestionSaferIn the danger zone
How high are the stakes attached to it?Watched, informative, low-consequenceWired to pay, budget, promotion, or a public ranking
How easy is it to move without doing the real work?Hard; the cheapest path is the work itselfEasy; an obvious shortcut moves it
Can you still see the real outcome directly?Yes, you spot-check the substanceNo, the number is all anyone looks at
Who is measured by it?People who also own the real outcomePeople rewarded only on the proxy
Is it one number or a balanced set?Several that trade off against each otherA single figure everyone chases

The pattern is that corruption needs high stakes and an easy shortcut and no independent view of the real thing. Remove any one and the law loses most of its grip. The cheapest fix is almost always the stakes: a number you watch instead of reward cannot corrupt the work much, because nobody has a reason to bend the work around it. The mistake teams make is attaching the heaviest consequence to the most convenient number, which is the exact combination Campbell's law punishes hardest. Before you tear out a metric that has already gone bad, though, it pays to find out why it was put there in the first place, because the pressure it created is often holding up something you cannot see.

Campbell's law at the scale of a field

The AI world is living through a textbook case, and it is Campbell rather than Goodhart because the damage is landing on a whole research process, not one system. When a benchmark becomes the high-stakes scoreboard for an entire field, labs optimize for the benchmark, because funding, press, and recruiting ride on topping it. Models get tuned to the benchmark's quirks, test questions leak into training data, and the score climbs while the capability it was meant to track moves less than the number suggests. The benchmark stops measuring the field's progress and starts measuring the field's skill at the benchmark, which is Campbell's distortion clause playing out across an industry. The single-model version of this, where one system learns to game its own reward signal, is the reward-hacking face of Goodhart's law; Campbell's contribution is the reminder that when the stakes are collective and high, the corruption is collective too.

The lesson transfers straight to how any team evaluates an AI tool. If you judge an assistant on one easy-to-score proxy, you will get a tool optimized for that proxy and a team that has quietly reorganized its own work to produce a good score, exactly as the deploy-frequency org did. The defenses are the same as for human metrics: keep the stakes on any single number modest, measure against a balanced set, hold back evaluations the system was not tuned to please, and keep looking at real output.

The metric whose reason got lost

Here is where Campbell's law turns from a bad quarter into a permanent institutional trap. Every high-stakes metric was chosen once, by someone, for a reason. Deploy frequency was elevated because releases used to be rare and terrifying. The response-time target was set because a specific customer churned after being ignored. At the moment of choosing, everyone knows the number is a proxy and roughly how much weight it can bear. Then the person who set it moves on, the original pain fades, and what remains is the number with its stakes still attached and its reasoning quietly decayed away. New team members inherit "we hit our deploy target" with no memory that the target was only ever a stand-in, and they optimize it with a clear conscience because nobody told them it was a proxy, or how high the stakes had crept.

That is the point of no return, and the fix is cheap and specific. When you attach real consequences to a metric, write down what it is a proxy for, how much weight it is meant to carry, and what would tell you it had started corrupting the work. That note belongs in the same durable record as the decisions it came from, next to the reasoning, so the next person inherits the purpose and the expiry conditions along with the number. A metric whose reason and stakes are written down can be dialed back the day it starts distorting things. A metric whose reason has decayed becomes a target no one dares lower and no one can defend. Try Scribelet free and keep the why and the intended weight behind each number, not just the number, where the people who inherit your dashboards will actually find it.

Reading Campbell's law correctly

The tempting overreaction is to conclude that measuring people is hopeless and to stop. That is the wrong lesson and an expensive one, because numbers are how you see a system too large to watch by hand, and steering blind is worse than steering on an imperfect gauge. Campbell's law is not an argument against measurement. It is an argument against putting life-altering stakes on a soft proxy and then looking only at the proxy. The skill is to measure while keeping the consequences proportional to how gameable the number is, and while holding on to an independent view of the real outcome.

So the standing question Campbell's law hands you is not "is this number accurate?" but "how much is riding on this number, and what is that doing to the work it watches?" Keep the stakes modest on anything easy to game, keep the real outcome in view, and write down why each measure exists and how much weight it was meant to hold. Try Scribelet free and let it hold the reasoning behind your metrics, so the numbers your team steers by keep pointing at the work instead of slowly replacing it.

Share this article

We use cookies for analytics to improve your experience. Learn more