Ranked on the Inc. 5000 list of America's fastest-growing companies
Data & Artificial intelligence (AI) 14 min read

Tokens or Headcount? The Wrong Question About AI Costs

Smartphone calculator resting on printed budget charts beside a laptop and notebook

Last Updated: October 8, 2026

In April 2026, Uber’s CTO, Praveen Neppalli Naga, disclosed that the company had used up its entire 2026 budget for Claude Code, Anthropic’s AI coding assistant. It was four months into the year (Quartz, citing Business Insider).

Uber got there early. Most enterprises running AI agents are on the same path.

AI token costs climb when companies run expensive models on every task and never measure what each task costs. Cutting experienced people to pay that bill usually makes things worse. The fixes are practical, and most of them start with measurement.

Uber’s president and COO, Andrew Macdonald, called it a “head-exploding moment.” He told The Verge that Uber would have to start weighing token use and its cost against headcount. CEO Dara Khosrowshahi had already said on an earnings call that Uber was slowing hiring to offset its AI spending. According to Forbes, a typical engineer’s bill ran $150 to $250 a month, while the heaviest users ran $500 to $2,000 (Quartz).

What’s in this article: Why the bill grows · Should you cut staff? · Cost levers compared · How to bring costs down · Is AI worth it? · What to do next · Scadea services · FAQ

Why is the AI bill growing when token prices keep falling?

Each token costs less, but agents and long prompts use far more tokens per task, so total spend climbs even as unit prices drop.

AI token costs are what model providers charge for the text a model reads and writes, billed per token. Bain & Company found that token prices fell by about half during 2025, but usage grew faster. AI agents that plan, call tools and retry multiply the token count, so the monthly bill keeps rising.

AI providers bill by the token, a small chunk of text, often part of a word. You pay for the tokens you send in and the tokens the model sends back.

Prices are falling fast. Bain found that the cost of tokens fell by half between December 2024 and December 2025. Bain’s summary of what happens next: “Average cost per token is falling, but token use is rising faster” (Bain).

AI agents drive much of that growth. An agent is an AI system that works through a task in steps on its own. It plans, calls tools, checks its work and tries again. Bain notes that agents burn tokens on multistep reasoning, error correction and loading context, and that tokens per query keep climbing as agents take on harder work. Our guide to agentic AI for enterprise workflows covers how those systems are built.

Habits add to the bill. Sebastian Gierlinger, VP of AI and IT at Storyblok, told Fierce Network that without guidance, staff default to the most powerful model, even for routine tasks (Fierce Network).

Should you cut staff to pay for AI?

Usually not. Klarna leaned on AI for customer service, saw quality drop and started hiring people again, and in some functions AI already costs more than people.

In February 2024, Klarna claimed its AI could do the work of 700 customer service agents (Entrepreneur). By May 2025, CEO Sebastian Siemiatkowski told Bloomberg that cost had weighed too heavily in the decision and quality had suffered. Klarna began hiring human agents again so customers could always reach a person (Fortune).

Klarna still runs on AI. A spokesperson told Fortune the company is “very much still AI-first.” The lesson is narrower. Cost alone was the wrong measure for that job.

Bain’s numbers show why the math changes from one task to the next. In software engineering, token spending is only about 1% to 2% of the cost of headcount. But in some domains, Bain found that agent and token costs already exceed the cost of offshore staff (Bain).

So swapping people for tokens can raise costs. It also drains the knowledge the AI depends on: the exceptions, the client history and the reasons a process works as it does.

How do the ways to cut AI costs compare?

Cutting staff lowers payroll fastest and carries the most risk. Tracking spend, routing models and trimming data cut the bill with far less downside.

Approach What it saves Main risk or effort Evidence
Cut staff to fund AI Payroll Quality falls and know-how leaves with the people Klarna rehired human agents in 2025 (Fortune)
Track cost per task and team Waste you can now see, through budgets and caps Needs tooling and an owner for the numbers Uber’s overrun went unmeasured until the budget was gone (Quartz)
Right-size and route models Premium-model spend on routine work Each task needs testing on the cheaper model Companies are adding caps and moving to cheaper and open-source models (PYMNTS)
Trim the data sent to models Input tokens, often with better answers Cleanup work up front More content in means more tokens burned (Fierce Network)
Keep people on sign-off Costly errors and compliance failures Review time belongs in the cost model An agent can do the analysis, and a person still signs the attestation (Bain)

How do you bring AI costs down without losing your best people?

Measure cost per task, match each model to its job, trim the data you send, price a role both ways before automating, and keep people on judgment calls.

Track spending by task and team first

You can’t cut what you can’t see. A single monthly invoice tells you the total and nothing about where it went.

Observability tools fill that gap. They show how many tokens each team, app and task uses, and what that costs. Once you can see spend per task, you can set budgets, watch for spikes and cap usage before the invoice arrives.

Macdonald made a related point at Uber: engineers who never see the invoice can treat AI tools as if they were free (Quartz).

Right-size the models

Right-sizing means matching the size and price of the model to the job. A top-tier model is worth paying for on a complex legal review. On sorting emails, it’s wasted money.

PYMNTS reports that companies are already moving this way. They’re adding usage caps, steering staff to the right tool for each task, switching to older and cheaper models, and adopting open-source models, which are openly published models a company can run on its own infrastructure (PYMNTS).

Many teams set up routing, so simple requests go to a small, cheap model and only hard ones reach the expensive model.

Run a “day in the life” cost check before automating a role

Before you automate any role, price out a normal day of that work both ways.

  1. List the tasks in a typical day for the role.
  2. Estimate the token cost of doing each task with AI, using real usage data where you have it.
  3. Add the human time still needed to review, correct and approve the AI’s output.
  4. Compare the total with the cost of the person doing the same tasks.

The result often points to a split. AI takes the repetitive tasks, and the person keeps the ones that need judgment, with more time to spend on them.

Clean up the data you feed AI

Every page of content you send a model costs tokens. Gierlinger’s warning: “the more tokens you will burn by ingesting this content” (Fierce Network).

Duplicate files, outdated documents and whole folders sent “just in case” all add cost. Trimming that data cuts spend and usually improves the answers too. Data quality pipelines catch much of it before it reaches a model.

Keep people where judgment or sign-off is required

Bain makes the point plainly: an agent that can do a compliance analyst’s work still needs a human to sign the attestation (Bain).

For banks, insurers and healthcare providers, that’s a regulatory requirement as well as good practice. Human in the loop means a person reviews or approves the AI’s output at set points. Plan for those points from the start, as described in human-in-the-loop design for regulated AI, and count that review time in your cost model.

Start with early wins you can measure

Pick a first project with a clear baseline: how long the task takes today, how often it goes wrong and what it costs. Agree on what success looks like before you start, so you’ll know within weeks whether it’s paying off. Our piece on measuring automation ROI beyond cost savings lists the measures worth tracking.

Choose the right platform for each job

Google Gemini, Anthropic’s Claude, OpenAI’s models and Microsoft’s Copilot tools each have different strengths and price points, and both change often. Test the options on your own tasks. Avoid sending every job to one platform out of habit.

Is AI still worth the cost?

Yes, when it’s measured and matched to the task. Bain expects the winners to track cost per task, right-size their models and redesign how they operate.

Bain’s view is that companies which manage token economics actively will come out ahead (Bain).

The risk of skipping that work is real. In June 2025, Gartner predicted that over 40% of agentic AI projects will be canceled by the end of 2027, pointing to rising costs, unclear business value and inadequate risk controls (Gartner).

Uber’s budget ran out in four months because usage grew faster than anyone measured it. The companies that come out ahead will keep their experienced people, give them AI that fits each task, and know what every task costs.

What to do next

Start with visibility. This month, get token spend broken out by team and by task, even roughly. Then pick the one role where someone has proposed replacing people with AI and run the day-in-the-life comparison on it, review time included. Those two numbers will tell you more about your AI budget than any vendor price list.

Scadea services for AI cost control

Scadea’s embedded specialists work inside client teams, next to the people who know the business, which is where AI cost control gets done. With 300+ consultants across 8 countries, Scadea works with enterprises in banking, insurance, healthcare and manufacturing on the groundwork that keeps AI spend predictable. An AI readiness assessment is the usual starting point: which tasks are worth automating and whether the data behind them is ready.

Data governance and quality covers the cleanup that cuts wasted tokens. Where people stay in the loop for review or sign-off, human-in-the-loop design sets those checkpoints, and enterprise AI orchestration routes work across models and tools.

Frequently Asked Questions

What is an AI token?

A token is a small chunk of text, often part of a word, that a model reads or writes. Providers bill for the tokens you send in and the tokens the model sends back.

Why did Uber run out of its AI coding budget so fast?

Use of Claude Code grew faster than anyone measured. Uber spent its full 2026 budget for the tool by April, with the heaviest users running $500 to $2,000 a month against a typical $150 to $250, according to reporting compiled by Quartz.

Did Klarna stop using AI?

Klarna still uses AI heavily. It started hiring human customer service agents again in 2025 after quality suffered, and a spokesperson told Fortune the company is “very much still AI-first.”

How much do AI tokens cost compared with headcount?

It depends on the function. Bain found that in software engineering, token spending is about 1% to 2% of headcount cost, while in some domains agent and token costs already exceed the cost of offshore staff.

What is model routing?

Model routing sends each request to the cheapest model that can handle it. Simple requests go to a small model, and only hard ones reach the expensive one.

What does AI cost observability track?

It tracks tokens and spend by team, app and task. With that view you can set budgets, spot spikes and cap usage before the invoice arrives.

Will falling token prices fix the AI bill?

Falling prices alone have not kept bills down. Bain found token prices fell by about half from December 2024 to December 2025, and token use is rising faster than prices are falling.

How many agentic AI projects get canceled?

Gartner predicted in June 2025 that over 40% of agentic AI projects will be canceled by the end of 2027, pointing to rising costs, unclear business value and inadequate risk controls.

When does replacing a role with AI make sense?

When a normal day of that work costs less with AI, including the human time to review and correct its output. Run the comparison first, because many roles end up split between AI and a person.

Read next: Agentic AI for Enterprise: Architecture & Governance

Let's Build Together

Let's build your next success story together.