Quartz
Subscribe
Quartz
Subscribe
Edition
Business News
A.I.
Technology
Money & Markets
Leadership
Lifestyle
Latest

Get Quartz in your inbox

Free daily briefing on global business news.

Business News
AirlinesAutomobilesFoodPharmaceuticalsPolitics & GovernmentRetail & EcommerceSpace & AerospaceEarnings
Technology
A.I.ComputingConsumer TechSpace & AerospaceEarnings
Money & Markets
Economic IndicatorsMarketsPersonal FinanceEarnings
Lifestyle
Cars & BikesCollectingEntertainmentFood & Fine DiningHealth and FitnessReal EstateTravel
Quartz

Global business news for a smarter world

Topics

  • Business News
  • Money & Markets
  • Tech & Innovation
  • Generation A.I.
  • Lifestyle
  • Leadership

Products

  • Daily Brief
  • Weekly Digest
  • Member Benefits
  • Quartz Pro

Legal

  • Sitemap
  • About
  • Accessibility
  • Privacy
  • Terms of Service
  • Advertising

© 2026 Quartz Media, Inc. All rights reserved.

A.I.

A 50-year-old economics law explains why AI tokenmaxxing was always going to fail

Every productivity metric eventually gets gamed. Tokenmaxxing is just the latest example of a decades-old organizational trap

By Anthony Lopopolo·6 min read·Updated July 3, 2026
Add QZ to Google
A 50-year-old economics law explains why AI tokenmaxxing was always going to fail

Josh Edelson / AFP via Getty Images

In 1975, a British economist named Charles Goodhart scribbled a footnote during a Reserve Bank of Australia conference that would outlive most monetary policy debates: "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." He was talking about inflation targets. He could have been talking about token leaderboards at Meta $META.

The pattern Goodhart identified has a simple modern translation, popularized by anthropologist Marilyn Strathern in 1997: "When a measure becomes a target, it ceases to be a good measure." The saying is true when an engineer runs agents in circles, generates documentation no one reads, or asks a frontier model what to have for lunch just to get to the top of an internal AI-usage ranking. It also explains why this keeps happening, with different technologies, in different decades, to companies that should know better.

The tokenmaxxing era, in which companies pushed employees to consume as many AI tokens as possible, wasn't the first time a new technology produced a metric that looked like progress and measured something else entirely. The organizational failure mode it represents has been named, studied, and warned about for decades. That makes it worth understanding on its own terms.

The metric always looks reasonable at first

Psychologist Donald T. Campbell articulated a parallel principle around the same time as Goodhart. Campbell's law states that the "more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor." Where Goodhart focused on the collapse of statistical regularities, Campbell focused on what happens to the humans doing the measuring. Both arrived at the same conclusion: quantitative targets warp the behavior they claim to track.

The mechanism isn't mysterious. Researchers Jongwoon Choi, Gary Hecht, and William Tayler gave it a clinical name in their work on strategic performance measurement: surrogation. It happens when managers stop using a metric as an imperfect proxy for a strategic goal and start treating it as the goal itself. At Meta, token consumption stopped being a signal of AI adoption and became the objective.

The conditions that produce surrogation are well documented: the strategic objective is abstract, the metric tied to that objective is concrete and distinct, and employees accept, at least subconsciously, the substitution of the metric for the strategy. AI transformation is about as abstract as corporate objectives get. Token counts are about as concrete as metrics get. The substitution was almost automatic.

Every technology wave finds its vanity metric

Eric Ries, author of "The Lean Startup," drew a distinction between vanity metrics and actionable metrics more than 15 years ago. Vanity metrics look impressive but offer no guidance for decisions. Actionable metrics tie specific, repeatable actions to observed results. Token consumption, absent any link to business outcomes, fits the vanity definition precisely, telling companies how much AI they were using without telling them whether the usage mattered.

The precedents stretch back further than the startup era. In software engineering, lines of code served as the default productivity measure for decades despite repeated warnings that it rewarded verbosity over quality. Bill Gates said it's "like measuring progress on an airplane by how much it weighs." Edsger Dijkstra noted that he had never seen anyone measure a composer's productivity by the number of notes scribbled monthly, yet the analogous measure for programmers persisted. An NBER working paper tracking more than 100,000 GitHub developers found the pattern alive in the AI era: coding agents led to a 741% increase in lines of code but only a 20% rise in actual software releases.

Communication tools followed the same trajectory. Research by Rob Cross, Reb Rebele, and Adam Grant, published in the Harvard Business Review, found that collaborative activity consumed 80% or more of knowledge workers' time, a figure that had risen 50% over two decades. Microsoft $MSFT's 2025 Work Trend Index found that the average worker receives 117 emails and 153 Teams messages per weekday, with interruptions arriving every two minutes during core hours. Volume rose. Whether output rose with it was a separate question that volume metrics could not answer.

Why the trap keeps working

Economists Bengt Holmström and Paul Milgrom provided a formal explanation in their 1991 multitasking model. When an agent has multiple tasks but only some are easy to measure, incentive structures push effort toward the measurable tasks and away from the rest. The more an organization rewards token consumption, the less attention flows to the hard-to-measure work that tokens were supposed to enable: better products, faster decisions, fewer errors.

This is why Palantir $PLTR CEO Alex Karp's critique of tokenmaxxing landed on a real structural problem. Karp compared the behavior to addiction, arguing that employees "are just sitting there all day" consuming AI output without producing business results. Palantir COO Shyam Sankar put it more precisely: "More tokens means more slop."

The broader data supports their skepticism. A report from MIT's NANDA initiative found that 95% of generative AI pilot programs failed to deliver measurable business impact, despite an estimated $30 to $40 billion in enterprise spending. McKinsey's 2025 state of AI survey showed that while 88% of respondents said their organizations use AI in at least one business function, only 39% reported any enterprise-level earnings impact. Adoption metrics tell one story. Outcome metrics tell another.

Measurement convenience is the root cause

The recurring pattern has a common root: organizations measure what is easy to measure, not what matters. Tokens are countable. Queries are loggable. Tool adoption rates are reportable to a board. The business value of an AI-assisted decision is none of those things.

The research on surrogation suggests that three interventions can help: involving the people who implement strategy in formulating it, loosening the link between metrics and incentives, and using multiple metrics rather than one. Nicole Forsgren, GitHub's vice president of research and strategy, has made a parallel case in software development. The DevEx framework she co-authored argues that developer experience can't be reduced to a single dimension and that activity metrics "should never be used in isolation either to reward or to penalize developers." Forsgren told The Pragmatic Engineer that "one of the most common myths" about productivity is "the notion that productivity is all about developer activity, things such as lines of code or number of commits."

Goodhart's law will not stop applying because companies learn it exists. The next technology wave will produce its own consumption metric, and some companies will build leaderboards around it. The question is whether the interval between adoption and correction can shrink. Right now, the evidence suggests it takes a few months and some very large bills.

Daily Brief

The essential business news, delivered fresh every morning.

Join 500,000+ readers who start their day with Quartz.

By subscribing, you agree to our Terms of Service and Privacy Policy.

Related

A.I.What happens when 12,000 tech employees become millionaires overnight
AutosTesla revenue climbed 26% but profit fell short of Wall Street expectations
Business NewsAlphabet posted 24% revenue growth as Google Cloud surged 82% in the second quarter
A.I.San Francisco's housing market is overheating. Tech buyers will make the situation worse
A.I.History suggests San Francisco homeowners will win big from the AI gold rush