They Called It Tokenmaxxing. I Call It Goodhart’s Law With a Credit Card.
There is a concept in economics called Goodhart’s Law. The short version: when a measure becomes a target, it ceases to be a good measure. You’ve seen it in sales organizations where reps optimize for activity metrics instead of revenue. You’ve seen it in customer support teams where handle time gets measured and suddenly every call ends faster but nothing actually gets resolved.
In the first half of 2026, some of the largest technology companies in the world ran a live demonstration of Goodhart’s Law at a scale that should go in a textbook.
They called it tokenmaxxing. And here’s what happened.
Companies rolled out AI coding tools to their engineering teams and, wanting to show progress on AI adoption, started tracking token consumption as a proxy for productivity. Use more AI, burn more tokens, look more AI-native. Leaderboards went up. Meta built one internally they called “Claudeonomics,” tracking usage across more than 85,000 employees and ranking the top 250 power users with titles like “Session Immortal” and “Token Legend.” Amazon ran something similar. Microsoft had an internal token dashboard where, notably, VP-level executives and distinguished engineers were showing up in the top rankings despite rarely writing code. Salesforce set minimum monthly spend targets, with a Mac widget updating every 15 minutes to show each employee where they stood against the floor.

What happened next should surprise no one who has ever watched a metric get gamed.
Engineers started asking AI to answer questions already covered in documentation, because doing so burned tokens while doing a manual lookup did not. They prompted agents to prototype features they had no intention of shipping, just to run up the count. They defaulted to agentic workflows on tasks they could have done faster by hand, watched the agent fail repeatedly, and let it loop through corrections because the meter was running and that was the point. One Microsoft engineer, worried about being flagged as insufficiently AI-native, described deliberately wasting tokens just to stay off a watch list.
At Salesforce, when engineers hit their monthly spending cap, the cap could be exceeded with a single button press. No approval required. Last week’s maximum became this week’s floor.
The Uber numbers tell the clearest story of where this led. According to reporting by The Information, Uber’s adoption of Claude Code jumped from 32 percent to 84 percent of a roughly 5,000-engineer organization between February and March 2026 alone. Average monthly spend per engineer ran between $150 and $250. Heavy users reached $2,000 per month. The company’s own CTO reportedly spent $1,200 in a single two-hour demo session. By April, Uber had burned through its entire planned 2026 AI coding budget. The year had barely started.
Uber’s chief operating officer then said publicly what a lot of finance teams were quietly discovering: there was no demonstrable link yet between all that token consumption and actually shipping better products.

The anonymous $500 million story made the rounds in late May. An AI consultant told Axios that one of their clients had spent half a billion dollars in a single month on Claude after failing to put any usage limits on employee access. Half a billion. In thirty days. The scale strains credulity, but the mechanism is straightforward: consumption-based pricing plus no guardrails plus a culture that rewarded usage equals a bill that finance discovered weeks after the fact.
A survey by cost governance firm Mavvrik covering 372 enterprises found that only 15 percent of companies forecasted AI costs within 10 percent of actual spend. A majority missed by 11 to 25 percent. Nearly one in four missed by more than 50 percent. The firm’s CEO had predicted the reckoning would arrive in the first half of 2026 as pilots flipped to production. He was right on schedule.
Microsoft, meanwhile, began winding down most internal Claude Code licenses across its Experiences and Devices division in June, with access ending June 30. The stated reason was toolchain consolidation around GitHub Copilot CLI. The timing was the last day of Microsoft’s fiscal year. Six months after the pilot launched, the company that has publicly said AI writes 20 to 30 percent of its code in some repositories decided it had seen enough of that particular experiment.

There is a Reddit comment from this period that deserves to be remembered. Someone pointed out that per-seat pricing was built for autocomplete, where a human typing is the natural cap on usage. Agentic tools removed that cap. One developer can kick off a task that makes hundreds of model calls on its own. The pricing model assumed humans were the bottleneck. That stopped being true.
That observation cuts to the actual problem underneath all of this. Tokenmaxxing was the symptom. The disease was that organizations deployed a new cost structure before they understood it, then layered in incentives that actively rewarded burning that cost as fast as possible, then discovered the bill weeks after the fact because governance frameworks were nowhere near ready for consumption-based AI pricing.
The Pragmatic Engineer’s Gergely Orosz, who covered the tokenmaxxing phenomenon in detail after speaking with engineers at Meta, Microsoft, and Salesforce, framed it cleanly: the industry was using token count numbers the same way lines-of-code metrics were used years ago. Easy to measure, easy to game, and almost entirely disconnected from the actual value being created.
Shopify handled this better than most. They built a usage dashboard early, renamed it away from “leaderboard” when the competitive framing became obvious, added circuit breakers that cut off access when personal spend spiked unexpectedly, and had leadership check in directly with the highest spenders to understand what they were actually doing. Their CTO told The Pragmatic Engineer that the more interesting question was not who spent the most, but whose tokens cost the most. Those turned out to be the engineers doing the most interesting work.
That is a policy built by people who understood what they were actually trying to measure.

The companies that built leaderboards and minimum spend floors were not staffed by foolish people. They were staffed by smart people who were under enormous pressure to demonstrate AI adoption, and who reached for the most available proxy for progress. Token consumption was visible, measurable, and easy to track. The fact that it measured nothing useful about outcomes was a secondary problem, and secondary problems have a way of becoming very expensive primary problems.
The AFP put it well in a June 2026 report: companies were in an era of “subsidized intelligence,” where investors were essentially footing the bill to allow AI to be offered cheaply enough that adoption would accelerate. The bill for that subsidy is now arriving, and it is arriving at the same time that the usage leaderboards are coming down and the finance teams are asking what exactly all those tokens produced.
Part 3 covers what happens next, why the overcorrection is already underway, and why the firms that were lean by necessity are now positioned to win.

This is Part 2 of a three-part series. Part 1 covered how large companies structurally prevent themselves from staying current after an acquisition, and why that pattern is repeating in AI. Part 3 covers the overcorrection already underway and why lean firms are positioned to win the next phase.
Sources: The Information (Uber Claude Code budget and adoption figures), The Pragmatic Engineer / Gergely Orosz (Meta Claudeonomics, Microsoft and Salesforce tokenmaxxing details), Axios / Madison Mills ($500M Claude spend report), Forbes / Janakiram MSV (Microsoft Claude Code wind-down, GitHub Copilot shift), Mavvrik/Benchmarkit enterprise AI cost survey (372 companies, forecasting accuracy data), AFP / Yahoo Finance (subsidized intelligence framing, Meta and Uber pullback), Gizmodo / AJ Dellinger (Meta token leaderboard, Amazon rollback).
Photos by [photographer names] on Unsplash.