← Back to Library

TokenBudgeting: Our conversations with enterprises on token spend

This piece cuts through the noise of sensational headlines about corporate AI spending sprees to reveal a far more nuanced reality: the era of unchecked "tokenmaxxing" was less a market-wide frenzy and more a series of isolated incentive failures. Dylan Patel brings on-the-ground data from over 50 enterprise conversations, challenging the narrative that major companies are burning through budgets at an unsustainable rate.

The Myth of the Spending Frenzy

Patel immediately dismantles the prevailing media story. "Widely reported responses to Tokenmaxxing budgets from companies like Meta and Uber are overstated and stem from poor incentives and employee allocation we didn't find present at other organizations," he writes. This is a crucial distinction. The article suggests that while headlines focused on specific outliers, the broader enterprise market has quietly shifted toward discipline rather than panic.

TokenBudgeting: Our conversations with enterprises on token spend

The author highlights how early adoption was driven by strange internal competitions. At Meta, an employee built a dashboard ranking power users, where one individual consumed roughly 280 billion tokens in a month just to climb a leaderboard. "Employees started competing for rankings like 'Token Legend' and 'Cache Wizard' by having agents do research for hours simply to burn tokens," Patel notes. This behavior mirrors the economic concept of Goodhart's law—once a measure becomes a target, it ceases to be a good measure. The dashboard was shut down within days, but the media narrative of unbridled excess persisted long after the experiment ended.

The headlines on early 2026 tokenmaxxing were a result of poor incentives and lax oversight rather than an absence of high ROI activities.

Patel argues that these stories are outliers. His data shows that even Meta, often cited as the most profligate spender, represents only a "3-5% customer" for major AI labs like Anthropic. The distribution of spending is heavily skewed: 90th percentile customers spend around $7,300 per employee annually, while the median spends just $136. This suggests that the market is not collapsing under weight but is instead maturing into a tiered structure where heavy usage is concentrated in specific, high-value roles.

The Reality of Budgeting and Conservation

As companies move from experimentation to production, the approach has shifted to strict budgeting, though there is no consensus on the numbers. Patel observes that budgets range wildly "starting at $250 and going up to tens of thousands a month." This lack of standardization reflects the early stage of AI integration, where organizations are still calibrating their tools.

Some companies have taken aggressive steps to curb costs. Uber, for instance, imposed a "$1,500/month/employee limit" after burning through its annual budget in just four months. In contrast, a top aerospace manufacturer capped employees at $250 a month. "Management believes that handing employees larger token budgets would push them to automate tasks that shouldn't be automated at all, like writing emails," Patel writes regarding one conservative firm's strategy.

This anti-automation stance is where the argument faces its sharpest friction. While management fears waste, Patel contends this view is short-sighted. "We believe this anti-automation view by management teams is naive," he asserts. He argues that email and communication workflows will inevitably become AI-native, and restricting tokens now may stifle long-term productivity gains. Critics might note, however, that in a recessionary environment, the immediate pressure to cut costs often overrides long-term strategic bets on automation.

To stretch their allowances, employees are getting creative. Patel points out that workers with Microsoft 365 subscriptions can "game the system by using Copilot's 365 chat to draft and synthesize ideas first, before spending metered tokens on Claude or Codex." This behavior highlights a gap in how companies track usage; if the tool isn't metered, it becomes an infinite resource, creating its own form of inefficiency.

AI as a Headcount Lever

Perhaps the most insightful reframing in the piece is how successful companies view AI spend not as an expense to be minimized, but as a lever for output. "Output expectations rise to match spend, and many workers have found themselves putting in even longer hours than before," Patel explains. The goal isn't just faster work; it's more work.

The article cites Amazon as a prime example of this dynamic. Despite public narratives about layoffs, the company is hiring at a faster pace because AI tools have unlocked efficiencies that allow for expansion. "A recruiter at Amazon responsible for scouting and placing principal engineers... noted that the process from initial screening call to team placement used to take 6-9 months, but with AI tools... that timeline been cut in half," Patel reports.

This shift is transforming how budgets are allocated. At a major US airline, token usage is now tied directly to project revenue, treating AI spend like travel or contractor fees. "When a project comes in, the financing team decides what percentage of revenue is set aside for expenses... and now that same budget must cover token usage as well," Patel writes. This financial discipline suggests that the market is moving past the hype cycle into a phase of rigorous ROI calculation.

Bottom Line

Dylan Patel's analysis provides a necessary corrective to the alarmist narrative surrounding enterprise AI spending, grounding the discussion in hard data from actual users rather than press releases. The strongest part of this argument is its demonstration that high spend is concentrated and productive, not wasteful and diffuse. However, the piece may underestimate the cultural friction of imposing strict caps on tools that employees view as essential to their daily efficiency. As organizations continue to refine these budgets, the real story will be whether they prioritize cost-cutting or productivity scaling—a choice that will define the next phase of AI adoption.

Deep Dives

Explore these related deep dives:

  • Goodhart's law

    The article describes how internal leaderboards for token usage caused employees to game the system by generating useless data, a textbook case of this economic principle where a measure becomes a target and ceases to be a good measure.

  • Prompt injection

    Understanding this specific security vulnerability explains why enterprises are now imposing strict budget caps and turning off premium tiers, as uncontrolled agent behavior can lead to runaway costs through malicious or accidental prompt manipulation.

  • Window (computing)

    The article's discussion of 'downgrading default models' hinges on the technical trade-off between a model's context window size and its token cost, which determines how much historical data an AI agent can process before hitting budget limits.

Sources

TokenBudgeting: Our conversations with enterprises on token spend

by Dylan Patel · SemiAnalysis · Read full article

It’s been reported that token consumption inside of enterprises is hitting a budgeting wall after unhinged consumption earlier this year. The SemiAnalysis team talked with over 50 customers by slack, phone, and at the Databricks AI Summit to understand trends within the enterprise.

Widely reported responses to Tokenmaxxing budgets from companies like Meta and Uber are overstated and stem from poor incentives and employee allocation we didn’t find present at other organizations

Budgets are now the new norm, but there’s no consensus number with budgets starting at $250 and going up to tens of thousands a month.

Companies are downgrading default models and turning off premium tiers while employees’ game subscription M365 Copilot usage to stretch their token allowance.

Rise and Fall of Tokenmaxxing.

Tokenmaxxing started earlier this year when companies like Meta and Salesforce began encouraging their employees to consume as many AI tokens as possible to boost productivity. At Meta, an employee even built a “Claudeconomics” dashboard that ranked the top 250 power users in the company. The results showed that Meta employees consumed over 60T tokens over a 30-day period, with the single highest individual accounting for roughly 280B tokens. Employees started competing for rankings like “Token Legend” and “Cache Wizard” by having agents do research for hours simply to burn tokens.

The dashboard was shut down 2 days later after The Information reported the spend.

That episode was just one amongst others in the enterprise tokenmaxxing trend in 1H26. Companies are now shifting focus from tokenmaxxing to token budgeting. Most recently, Uber made headlines for burning through their Claude Code and Codex annual budget in four months. In response, the company imposed a $1,500/month/employee limit, with over-limit requests allowed and approved on a case-by-case basis. To see if the news reports on early 2026 tokenmaxxing and now tokenbudgeting were true, the SemiAnalysis team conducted on-the-ground conversations at the Databricks AI Summit and talked with large enterprises to understand the trends.

Our View of the Data & Narrative.

There are many news stories out there on tokenmaxxing and resulting budget blowouts. However, in our work in the Tokenomics Model, we estimate that 90th+ percentile customers make up most of the revenue and are at very little risk to API revenue cuts through the rest of the year. Even Meta, who was burning through 70T tokens per month in February and is spending close to at ...