August 25, 2026

Inside the AI Spend Management Playbooks at AT&T and Schneider

Token prices are falling, yet enterprise AI bills keep climbing. Learn how AT&T, Schneider Electric, and Uber brought AI spend under control.

10 min read

  • Uber burned its entire 2026 AI coding budget in just four months.

  • AT&T’s AI spend management practice cut costs up to 90% by routing work to the right model.

  • Falling cost per token drives higher usage, so total AI spend still climbs.

Staff writer

From AI to FinOps, our team's collective brainpower fuels this blog.

In April 2026, Uber's chief technology officer did the math and was unpleasantly surprised.

The company had burned through its entire 2026 budget for AI coding tools … in four months. It was meant to last the year, but by April the line item was already in the red.

The tools were Anthropic's Claude Code and Cursor, handed to roughly 5,000 engineers in late 2025. Usage climbed fast, helped along by an internal leaderboard that ranked teams by how much AI they used. By spring, some engineers were running up monthly bills between $500 and $2,000. CTO Praveen Neppalli Naga told The Information he was "back to the drawing board" on the company's assumptions about AI spend management. Uber's total research and development spend was $3.4 billion in 2025, so these were tools well within its price range. Yet finance teams had never budgeted for the pricing model.

Uber was not alone in turning usage into a game. At Meta, engineers built an internal leaderboard called "Claudeonomics" that ranked the company's top AI users by how many tokens they consumed. In one 30-day stretch, employees ran through more than 60 trillion tokens, a figure estimated at around $9 billion at public prices, with the heaviest single user averaging 281 billion tokens. Leaderboards like these did exactly what they were designed to do, driving adoption through the roof. A critical question was left out of the mix: what is that volume of consumption costing … and what is it earning?

With AI token usage amongst AI agents alone expected to multiply 24x in less than four years, the question is more pressing than ever.

Token usage by AI agents is expected to multiply 24x by 2030, furthering the importance of AI spend management. A line chart from Goldman Sachs shows token use is expected to reach close to 120 quadrillion monthly tokens for enterprise agents in 2030.
Source: Goldman Sachs — Token usage by AI agents is expected to multiply 24x by 2030, furthering the importance of AI spend management.


The companies prioritizing AI spend management plan well in advance. For example, NVIDIA CEO Jensen Huang has said he fully expects a $500,000 engineer to consume at least $250,000 in tokens. Because of that foresight, these companies build cost discipline (budget caps, model routing, and outcome-based accountability) into the operating model before scaling, not after the invoice arrives.

What Is the Tokenomics Paradox?

Uber has thousands of engineers and one of the more sophisticated technology organizations in the business. If Uber can lose the thread on AI spend this fast, other organizations should perk their ears.

A simple line graph shows the cost of tokens decreasing over time while AI usage increases, creating an “X” shape. As bills continue to rise, AI spend management becomes all the more essential.
The tokenomics paradox at work: cost per token has decreased over time, yet AI usage has increased. As bills continue to rise, AI spend management becomes all the more essential.

The answer starts with a paradox that sits at the heart of AI tokenomics. The price of a single AI token keeps falling, yet the total AI bill keeps climbing. Recent MIT research found that the cost of frontier-level performance is dropping by up to 10x a year. Yet enterprise AI spend is surging in response: cheaper tokens invite far more use, and teams reach for the most expensive models by default. 

A strategic decision hides beneath the paradox, and it is the one most leaders haven't made on purpose. A wave of capable open-weight models, many of them from Chinese labs like DeepSeek, Alibaba's Qwen, and Moonshot's Kimi, now runs at a fraction of the price of closed frontier models from the big American providers. The pull is already showing up in real traffic. Through the first half of 2026, U.S. companies routed a rising share of their tokens to open models, reaching as high as 46% in some weeks, up from an average near 11% a year earlier.

A line graph shows Chinese AI model usage by U.S. companies spiking significantly through the summer of 2026, with data beginning in January of 2025. U.S. company share of Chinese model tokens reached 44.99 in June of 2026.
Source: CNBC / OpenRouter — Chinese AI model usage by U.S. companies has risen significantly during 2026, according to OpenRouter. Organizations are increasingly turning to model choice as part of their AI spend management strategy.

Why Does AI Spend Run Out of Control?

The cost model for AI is not the cost model enterprises planned around. Traditional software bills by the seat, a flat and predictable number you multiply by headcount. AI bills by consumption, which moves with every prompt, every agent, and every choice about which model does the work. That shift is what makes AI spend management a new discipline rather than a line on an old budget, and it creates three distinct problems:

  1. Paying premium prices for routine work: when one general-purpose frontier model handles every task, the simplest jobs carry the same high AI token cost as the hardest ones.
  2. AI Spend that no one can tie to business value: without a clear link between a dollar spent and an outcome earned, money flows toward activity instead of results.
  3. Adoption with no ceiling: consumption-based pricing has no natural stopping point, so usage can outrun the budget before anyone notices.

Here’s how three companies handled these issues.

How AT&T Brought Down Its AI Cost Per Token

AT&T ran straight into the first problem at a scale few companies reach. Its internal assistant, Ask AT&T, was processing 8 billion tokens a day. Pushing all of that through large reasoning models was neither fast nor affordable, and Chief Data Officer Andy Markus knew it. The company needed a way to meet that demand without paying frontier prices for every request, including routine ones that didn't need a frontier model at all.

The strategy question was simple to state and hard to answer: which model should handle which task? AT&T's answer was to stop wiring its applications to any single model. The company built its stack around models that are, in Markus's words, "interchangeable and selectable," and it made a point of "never rebuilding a commodity." 

A routing layer sends each employee request to the cheapest model that can handle it, reserving the expensive frontier models for work that genuinely needs them. The company folded in smaller open-weight models alongside the premium ones, so routine jobs stopped drawing premium prices. Because the market shifts week to week, the architecture is built to swap components as better or cheaper options emerge.

The results landed on both sides of the ledger. For some high-end work like code generation, costs fell by as much as 56% while quality slipped only about 2%. AT&T rebuilt the orchestration layer around small language models, large "super agents" directing smaller "worker" agents that each do concise, purpose-built jobs.  Markus's team cut AI costs by up to 90% even as daily volume more than tripled, from 8 billion tokens a day to 27 billion.

"I believe the future of agentic AI is many, many, many small language models. We find small language models to be just about as accurate, if not as accurate, as a large language model on a given domain area."

Andy Markus, Chief Data Officer, AT&T

The choice of model is no longer left to habit. AT&T runs rigorous evaluations on outside options and its own tools, keeps a human in the loop overseeing the chain of agents, and enforces logging, data isolation, and role-based access as agents hand work to one another. Markus's guiding warning is against overbuilding: teams should ask whether a job even needs to be agentic, or whether a simpler, single-turn solution would be cheaper and more accurate. Accuracy, cost, and responsiveness are the three principles his team returns to.

The lesson for a CIO is that the model is a cost lever, not a default: a company that can switch models can manage its bill.

Schneider Electric Ties Every AI Dollar to an Outcome

The loudest question in enterprise AI right now is whether the spend is producing anything. At Uber, the company’s president and chief operating officer, Andrew Macdonald, admitted the company could not yet draw a line from its rising AI use to features customers actually feel. "That link is not there yet," he admitted, adding that it was very hard to connect internal AI stats to something like shipping 25% more useful features.

Schneider Electric, the French industrial and energy-management company, built its whole approach to avoid that trap. Many companies fall into what amounts to technology tourism, running thousands of experiments to see what sticks. Schneider went the other way. It requires every AI initiative to prove clear business value and to plan for full-scale deployment before it moves forward. Chief AI Officer Philippe Rambach has teams carry a use case from first idea through production. A promising demo must lead somewhere.

The payoff is a portfolio that earns its cost. Schneider now runs close to 100 AI use cases in production, split between customer-facing and internal operations. Its AI handles 7.5 million customer service tickets a year. A self-healing supply chain built on these tools has driven a 10% drop in inventory, a 15% average improvement in production yield on targeted lines, and more than 100 million euros in value.

"We always start from the business and customer needs, pain points of employees, where AI can help."

—Philippe Rambach, Chief AI Officer, Schneider Electric

Starting from the business need, not the technology, controls costs. A use case that must prove its outcome before it scales cannot turn into an untracked bill. And the discipline surfaces problems money alone would hide. One pilot meant to predict which sales bids Schneider would win ended up exposing a flaw in the underlying data that the business needed to fix first, exactly the kind of finding a value-first review is built to catch. 

The approach is slower at the start than turning everyone loose with a frontier model. But AI spend management becomes clear by the time a use case reaches production: it can’t get there without a reason.

How Uber Could Turn a Blowout Into AI Spend Management

Uber's story closes this framework, because its budget blowout is the clearest picture of the third problem: adoption with no ceiling. 

The budget vanished for structural reasons. Claude Code does not price by the seat. It meters tokens, so the same engineer on the same day can generate wildly different invoices. An annual budget built around predictable per-seat costs cannot absorb that kind of swing. This is the real break from traditional software: consumption pricing moves with behavior, and behavior, especially when a leaderboard is egging it on, can be inconsistent.

Uber's first fix was blunt but necessary. The company set a $1,500 monthly cap per employee for each agentic coding tool, tracked through an internal dashboard every employee can see, with room to exceed the cap for a real reason.

What the discipline requires is a new operating model for consumption-based spend, and consulting firm EY has laid out one of the clearest versions. EY calls the role an "Agent FinOps Lead" and argues that agentic AI shifts enterprise cost from fixed software and labor to variable compute, a kind of spending finance teams have never had to manage before. 

Its framework centers on a few moves. Give every stream of AI spend a clear owner. Benchmark cost on a per-task or per-outcome basis, so you know what a unit of work should cost and can tell when an agent is running away. Then install hard controls, spend ceilings, call-volume caps, and automatic shutoffs at the level of the individual agent, the workflow, and the business unit. "Without circuit breakers," EY warns, "the bill is only visible after the damage is done."

That last line is the whole lesson of Uber’s budget overspend. The company gave thousands of engineers powerful agents before it decided what those agents were allowed to cost. The governance question is not how to slow AI down, but how to give consumption a shape, an owner, and a limit before the invoice arrives, not after. A cap tells an engineer when to stop. An AI spend management model tells the business what the spend was for.

What to Take Into Your Next AI Tokens Budget Review

AI bills like electricity, not software, and the companies staying ahead are the ones metering it on purpose.

The pattern across AT&T, Schneider Electric, and Uber is that cheaper models do not produce a cheaper AI bill on their own. AI spend management does. AT&T built the ability to route work to the right model. Schneider made every use case earn its cost before scaling. Uber learned, the hard way, that consumption needs a ceiling and an owner from day one.

So before you approve the next increase in AI spend, bring three questions to the table:

  1. Which of our use cases can name the specific business outcome they are paying for? 
  2. Which of our agents and tools have a real cost ceiling and a named owner? 
  3. Can we route work to a cheaper model when the task allows, or are we locked to one provider and paying premium prices for routine work? 

The companies that can answer those questions, and treat AI tokenomics as a priority, are the ones who will still recognize their AI bill a year from now.

Cut through the AI hype and join the thousands of business leaders getting practical enterprise insights delivered to their inbox

Welcome to the community! We'll be in touch soon.

Frequent Asked Questions

What is Agent FinOps?

+

It is an operating model for consumption-based AI spend, as described by EY. It gives each spend stream an owner, benchmarks cost per task or outcome, and installs circuit breakers like spend ceilings and shutoffs at the agent, workflow, and business-unit level.

How does Schneider Electric tie AI spend to value?

+

Every AI initiative must prove clear business value and plan for scale before it advances. One team carries a use case from idea to production. The result: nearly 100 use cases live, 7.5 million tickets handled yearly, and over 100 million euros in value.

How did AT&T reduce its AI costs?

+

AT&T stopped wiring apps to one model. A routing layer sends each request to the cheapest capable model, reserving frontier models for hard work. Code-generation costs fell up to 56% with only a 2% quality drop, and overall AI costs fell up to 90% as volume tripled.

Why do AI bills rise when token prices are falling?

+

Cheaper tokens invite far more usage, and teams default to premium models. When volume grows faster than the per-token price drops, total spend climbs even as each token costs less. MIT research found frontier performance costs falling up to 10 times a year.

What is AI spend management?

+

AI spend management is the practice of governing consumption-based AI costs. Unlike per-seat software, AI bills by tokens used, so it needs owners, cost benchmarks per task, and hard limits like spend ceilings and automatic shutoffs to keep total spend tied to business value.