The Overcorrection Is Coming. Lean Firms Won’t Wait for It.
Here is what always happens after a technology binge.
The early adopters go all in, often without governance frameworks to match their ambition. The bills arrive. Leadership overcorrects. Access gets restricted, licenses get canceled, leaderboards come down, and the finance team gets a seat at the table it probably should have had from the beginning. The technology doesn’t go away. But the organization swings from unconstrained enthusiasm to cautious rationing, and somewhere in that swing, the firms that were never in the binge phase quietly pull ahead.
We are in the early stages of that swing right now.
Microsoft began canceling Claude Code licenses across its Experiences and Devices division in June 2026, just six months after the pilot launched. Uber’s leadership publicly acknowledged that burning through an entire annual AI budget in four months produced no demonstrable improvement in shipping velocity. Meta shut down its token leaderboard after internal details leaked showing the top “Token Legend” had burned through 281 billion tokens in a month, more compute than reproducing the entirety of Wikipedia 33 times over. Amazon followed, axing its AI usage scoreboard after it became clear employees were gaming it with pointless tasks just to hold their position.
The AFP reported in June 2026 that the era of “subsidized intelligence” is ending. The AI companies that priced their tools cheaply to accelerate adoption are preparing for IPOs and need to demonstrate a path to profitability. Anthropic filed confidentially with the SEC. OpenAI is expected to follow. When public market investors start asking about unit economics, the subsidy that made tokens feel nearly free starts getting priced more honestly.
A Gartner report cited in late 2026 coverage predicts that inference costs for generative AI models in 2030 will be roughly a tenth of what they were in 2025. That sounds like good news, and eventually it will be. But the same report projects token usage growing five to thirty times current levels, largely driven by agentic workflows that consume orders of magnitude more compute than simple chat queries. One agent-driven task can burn a thousand times the tokens of a standard prompt. The cost per token may fall. The total bill may not.
That math lands differently depending on whether your organization spent the last eighteen months building discipline around AI use or performing it.
Large organizations are now doing what large organizations do after an expensive mistake. They are implementing controls, tiering access, restricting premium models to approved use cases, and demanding ROI metrics before approving new deployments. Some of this is appropriate and overdue. Some of it will overshoot. A blanket restriction that locks a genuinely productive engineering team off the tools that were actually working is a different kind of expensive, just slower and quieter.
The firms that will navigate this best are not the ones scrambling to install guardrails after the fact. They are the ones that never had the budget to run without guardrails in the first place.
I have written before about why startups eat large companies’ lunch on technology. The argument I made then was about vendor lock-in policies preventing large firms from using differentiated features. The argument here is similar but the mechanism is different.
A startup with a $50,000 annual AI budget and three engineers learns very quickly what produces value and what burns compute for no return. They cannot afford the leaderboard phase. They cannot afford the agent loop that runs for six hours on a task that a junior developer could have done in forty minutes. Constraint forces the question that abundance never has to ask: is this actually working?
That forced discipline is an asset that does not show up on any balance sheet, but it compounds. The team that has been iterating on what genuinely works for eighteen months while the enterprise was busy tokenmaxxing has a lead that is not just technical. It is organizational. They know what the tool is good for. They know where it breaks. They have built workflows around its actual capabilities rather than its marketed ones.
When the enterprise finishes its overcorrection and comes back to AI with a more disciplined posture, it will find that the gap it thought it was closing by burning through budget in 2025 and 2026 has not closed at all. In some cases it will have widened.
This is the part of the pattern I watched play out in telecom in the early 2000s. The acquiring company’s procurement policy, designed to protect manufacturing scale, created a technical lag that compounded over years. The solution, eventually, was to write large checks to acquire whoever had stayed current. The acquisition premium was the deferred cost of the policy.
The AI version of this cycle is moving faster, because everything in software moves faster. But the structure is identical. Large organization imposes constraints that make sense locally but create lag systemically. Lean firms operate without those constraints, not because they are smarter but because they cannot afford them. The gap grows. The large organization acquires.
There is a version of this story where the large organizations learn from the tokenmaxxing episode and build genuinely disciplined AI practices that let them stay competitive. Shopify’s approach, which I covered in Part 2, is a reasonable template. Usage dashboards rather than leaderboards. Circuit breakers that catch runaway spend automatically. Leadership that checks in with high spenders to understand the use case rather than just reward the number. ROI metrics that tie spend to outcomes. That is a mature operating model and it is achievable.
But building it requires admitting that the last phase was not a success, and large organizations are not always fast at that admission. The finance team finding out about a six-month budget burn at the quarterly close is not a sign of an organization that is about to move quickly on governance reform.
The firms to watch are not the ones announcing the largest AI investments. They are the ones quietly shipping products that work, built by small teams who figured out how to use the tools without a leaderboard telling them to.
They have been doing this the whole time. They just were not making noise about it.
The acquirer, eventually, will find them. And it will pay a premium. It always does.
This is Part 3 of a three-part series. Part 1 covered how large companies structurally prevent themselves from staying current after an acquisition, drawing on a firsthand experience from the late 1990s telecom boom. Part 2 covered the tokenmaxxing episode, what it actually was, and what it cost the organizations that ran it.
Sources: AFP / Yahoo Finance (subsidized intelligence era ending, Anthropic IPO filing, Meta and Uber pullback); Forbes / Janakiram MSV (Microsoft Claude Code wind-down, GitHub Copilot transition to usage-based billing); Gizmodo / AJ Dellinger (Meta token leaderboard details, 281 billion token figure, Amazon rollback); The Pragmatic Engineer / Gergely Orosz (Shopify circuit breakers and usage dashboard model); Android Authority / Tushar Mehta (Gartner inference cost projection, token usage growth forecast); Stocktwits / Rounak Jain (tokenmaxxing decline, Goodhart’s Law framing in analyst commentary).
Photos by [photographer names] on Unsplash.