Why AI Power Users Are Quitting ‘Tokenmaxxing’ for Efficiency

Earlier this year, the fastest way to look productive at a lot of tech companies wasn’t shipping more features — it was burning more AI tokens. Internal leaderboards ranked engineers by how much Claude and Codex usage they racked up. One Meta employee reportedly burned through 281 billion tokens in a single month. The trend even had a name: tokenmaxxing. And by the middle of 2026, it had already collapsed under its own weight.

How “burn more tokens” became a status symbol

The logic behind tokenmaxxing was seductively simple: if AI usage correlates with productivity, then more usage must mean more productivity. Companies leaned into it hard — some formalized informal leaderboards tracking who used the most AI, treating token counts the way sales teams treat call volume. It’s the kind of metric that looks great in a quarterly slide until someone asks the follow-up question: burned tokens on what, exactly?

That follow-up question is what killed it. Uber blew through its entire 2026 AI token budget in roughly four months, with around 5,000 engineers pushing consumption well past projections, and the company’s COO Andrew Macdonald later admitted the team struggled to “draw a direct line” between all that spend and any useful feature or functionality actually shipped to users. When the CFO asks for that line and nobody can draw it, the leaderboard stops looking like a productivity metric and starts looking like a very expensive vanity number.

Supercomputer server room — AI compute cost
Photo: Wikimedia Commons (CC BY-SA)

The interesting tangent: Goodhart’s Law, live in production

Here’s the part that’s genuinely fascinating if you like watching incentive structures break in real time. Reports out of Amazon found employees quietly generating meaningless tasks specifically to inflate their own token counts once usage became something managers watched. That’s Goodhart’s Law playing out with almost comedic precision: “when a measure becomes a target, it ceases to be a good measure.” Nobody set out to create fake work — the system just rewarded a number, and people optimized for the number instead of the thing the number was supposed to represent.

Salesforce felt the financial end of this from a different angle. The company’s Anthropic bill reportedly reached around $300 million annually, and CEO Marc Benioff has been candid about the missing piece: no “smart router” to send each query to a model sized appropriately for the task. Paying frontier-model prices to summarize a routine internal email is the AI-era equivalent of hiring a surgeon to put on a Band-Aid — technically capable, wildly overpriced for the job.

The part nobody put on the leaderboard: burnout

The financial waste turned out to be the easy problem to see. Research from the DORA team — the group behind the industry-standard DevOps performance metrics — found darker consequences underneath the token leaderboards: technical debt piling up from code nobody fully reviewed, developers hoarding knowledge instead of collaborating because the metric rewarded individual output, and creeping job insecurity as “how much AI did you use today” started to feel less like a productivity nudge and more like a surveillance number. Teams were reportedly spending up to 10 times more for only marginal gains in actual throughput — the AI equivalent of redlining an engine that isn’t actually going any faster.

There’s also a trust problem hiding underneath all of this that rarely makes the headlines: DORA’s research puts the share of developers who still don’t fully trust AI-generated output at around 30%. That’s nearly a third of the people being measured on how much of a tool they use privately doubting whether the tool’s output is even good. Tokenmaxxing didn’t just measure the wrong thing — it measured the wrong thing while a meaningful chunk of the workforce quietly worked around the tool to protect their own output quality.

The reversal: from tokenmaxxing to tokenomics

By late June, the correction was well underway. Meta quietly removed its internal token-usage leaderboard. Microsoft canceled Claude Code subscriptions across several product divisions rather than keep paying for usage nobody could tie to output. Startups moved even faster and cheaper: Lindy cut its inference costs by roughly 90% simply by routing more of its workload to DeepSeek instead of defaulting every query to a frontier model.

That last example is the one worth sitting with. A 90% cost cut from model routing alone — not from using AI less, but from being deliberate about which model handles which task — is a bigger efficiency gain than most companies get from an entire year of “AI productivity initiatives.” It turns out the frontier model isn’t always the right tool; it’s just the default one, and defaults are exactly what get exposed when the budget runs out early.

What this actually reveals about the current AI moment

The uncomfortable truth underneath tokenmaxxing’s rise and fall is that most companies still don’t have a clean answer to “what is AI actually worth to us?” — so they reached for the easiest proxy available, the same way page views stood in for content quality in the early web, or lines of code stood in for engineering output in the 1980s. Roughly 95% of enterprise AI usage today still runs on frontier-tier models by default, according to industry estimates, which means there’s an enormous amount of routing and right-sizing savings still sitting on the table for companies that haven’t gone through their own Uber moment yet.

We’ve been tracking this same theme of enterprises discovering that AI dependency comes with sharper edges than expected — see our look at what a 19-day model shutdown taught businesses about treating AI as infrastructure. Tokenmaxxing and that story are really the same lesson wearing different clothes: businesses adopted frontier AI fast, on the assumption that more access and more usage were unambiguously good, and are only now building the operational discipline — cost controls, model routing, actual ROI measurement — that any other major line-item expense gets from day one.

If you’re not running an enterprise AI budget, this might still sound like someone else’s problem — but the same logic quietly applies to anyone paying for a premium AI subscription out of habit rather than need. The individual version of tokenmaxxing is defaulting to the most expensive model tier for every task, including the ones a cheaper or free tier would handle just as well. The enterprise correction and the personal one are the same insight at different scales: usage isn’t value, and “more AI” was never actually the goal. Shipping something useful was.

It’s also worth watching who benefits from the correction. Cheaper, leaner models suddenly look a lot more attractive once “use the biggest model for everything” stops being free money — which is exactly the pitch smaller labs like Moonshot AI have been making all year. Tokenmaxxing’s collapse isn’t really a story about AI getting worse. It’s a story about the free-spending phase of a new technology finally meeting a budget meeting — and losing.

Deixe um comentário

O seu endereço de e-mail não será publicado. Campos obrigatórios são marcados com *