Why Is AI So Expensive to Run? The Truth About AI Token Pricing

Introduction

Why is AI so expensive to run behind the scenes, even when you use it for free? Millions of people open ChatGPT, Claude, or Gemini every day without paying a cent. That feels like a great deal for users. However, behind every free response sits a real bill that someone eventually has to cover.

Microsoft, Google, and Anthropic have poured hundreds of billions of dollars into building large language models. Naturally, these companies want that investment back. So they created paid tiers with extra features for coding, billing, and business tasks. Meanwhile, a whole new wave of third-party firms is building AI agents on top of these models to handle specific jobs automatically. This spending race is also why Google’s AI cash flow has turned negative even as its products grow.

The problem is nobody has fully figured out how to price any of it yet. That matters for readers because this pricing chaos is already shaping the subscription costs, tool limits, and surprise charges people are starting to notice, and it even feeds into wider fears about whether AI will replace your job.

Why Is AI So Expensive to Run for Companies?

Why Is AI So Expensive to Run for Companies?

“Setting a fair price shows why AI is so expensive to run — it sounds simple on paper.”
In practice, it is turning into one of the trickiest problems in the tech industry today.

“Trying to tie someone into a cost model for the next 12 months, two years, three years, it doesn’t make any sense, honestly, because we don’t know,” says Simon Gooch at Saviynt, an identity management company incorporating agentic AI into its services.

The root of the issue is tokens, the mathematical building blocks that AI models use to process language. When someone types a prompt into a chatbot, that text gets broken into tokens. The model’s response also comes back as tokens, which then get converted into readable text, code, or commands.

This process is far from predictable. Small changes in a prompt can produce completely different answers. Additionally, the exact same prompt will not always generate the exact same response twice. Different models, such as Moonshot AI’s Kimi K3 or Qwen3 Max, handle the same request in different ways too.

How Agentic AI Makes Costs Even Harder to Predict

Things get messier with agentic systems, where businesses run multiple AI agents together to make decisions and take action. This setup increases both token consumption and unpredictability at the same time.

According to Goldman Sachs analysis, the cost of individual tokens has dropped sharply in recent years. Yet total token consumption by businesses and consumers has exploded regardless. The bank forecasts that token use will increase 24 times between 2026 and 2030, reaching 120 quadrillion tokens a month as companies shift toward AI agents, a shift already visible in efforts like Nokia’s AI-RAN platform built with Nvidia.

Consequently, many companies and individuals have only a rough idea of how many tokens they are burning through until they run out or receive a shocking monthly bill. Microsoft has reportedly pulled back its engineers’ use of certain third-party coding tools. Uber reportedly burned through an entire year’s AI coding token budget in just a few months this year.

How Are Companies Trying to Control Their AI Costs?

How Are Companies Trying to Control Their AI Costs?

Will Venters, Associate Professor of Digital Innovation and Information Systems at the London School of Economics, says businesses regularly get caught off guard while experimenting with or rolling out AI internally.

“People are finding it really hard to manage that cost… it’s a non-deterministic output, so it’s a non-deterministic value,” he said.

Some companies have found workarounds, at least for now. Oliver King-Smith, founder of engineering software firm smartR AI, notes that smaller organizations can fly under the radar using flat-fee personal accounts that big vendors likely dislike.

“This has to end at some point in time, because the big guys are taking a bath on those accounts,” King-Smith explains. Once shareholders start pushing for profits, he predicts the major AI platforms will begin clamping down on these workarounds. Google has already taken a similar step by freezing its AI chip design to improve Gemini’s efficiency rather than spending endlessly on new hardware.

King-Smith also suggests companies need to think more carefully about which AI models they actually use for which tasks.

Why Prompt Precision Matters for Cost Control

Precision matters just as much as model choice. Rob Steele, CFO at UK accounting software firm iplicit, compares vague AI prompts to sending someone shopping without a list.

“You wouldn’t send someone in your family out to get the weekly shop without any kind of detailed instructions as to what you expect in that shopping basket, right?” Steele says.

“This challenge grows significantly once a company builds AI into a product used by thousands of people. Venters points out that costs can balloon quickly…” Venters points out that costs can balloon quickly. Managers often realize too late that tokens are needed not just for core development work, but also for testing, security, and guardrails, concerns that echo cases like xAI being sued over Grok-generated sexual deepfakes of minors, where safety spending became unavoidable.

“It’s particularly hard when you’re looking at agentic processes,” Venters says. Adding more AI agents takes just a click of a button. Expanding a human workforce, on the other hand, requires careful planning around headcount and hiring.

Still, Venters offers a more optimistic angle too. Unpredictable token costs might mean companies are actually getting more value from their AI use, similar to how the NHS has used AI tools to cut waiting times despite the added expense.

“It’s not quite the same as a calculator,” he says. “The more you give it, the more expensive it is, but the better the result may be.”

Who Actually Pays for Expensive AI in the End

Who Actually Pays for Expensive AI in the End

Eventually, businesses have to pass these costs onto their own customers. Figuring out exactly how remains an open question across the industry.

“Nobody’s really figured it out,” says Bill Peterson, senior director of product marketing at Sumo Logic. His company is previewing new AI-powered security services and is still working out how to charge customers for them.

“We’re still having some fun conversations about this internally,” Peterson admits.

A few pricing models are on the table. Companies could simply raise prices across the board. Alternatively, they might charge based on results delivered, or sell bundles tied to specific incidents rather than raw usage.

However, whatever approach a company settles on could be upended the moment large language model providers change their own pricing strategies. “You get into variable pricing, and it’s changing every couple of months,” Peterson says. “Customers don’t like that. That’s not how anybody builds a budget.” Even regulators are watching AI costs closely, as seen in the UK threatening big tech with penalties over child safety features, which adds compliance spending on top of compute bills.

Conclusion

The question of why AI is so expensive to run comes down to one core issue: usage keeps climbing even as individual token prices fall. Businesses are stuck trying to budget for something that behaves nothing like a traditional software subscription.

As agentic AI adoption grows, this problem will only become more visible to everyday users too. Ultimately, the winners in this space may not be the companies with the best models, but the ones that figure out pricing structures customers can actually trust and plan around.

FAQs

Why is AI so expensive to run even when it feels free to use?

AI models process every prompt and response as tokens, and running the underlying compute for those tokens costs companies real money, even when users pay nothing at the point of use.

What are tokens in AI, and why do they matter for pricing?

Tokens are the mathematical chunks that AI models use to process prompts and generate responses. The number consumed directly affects how much a company or user ends up paying.

Why is it so hard for companies to predict their AI costs?

AI outputs are non-deterministic, meaning the same prompt can produce different answers each time, and different models respond differently to the same request. This makes forecasting token usage and budgets difficult.

How much will AI token consumption grow in the coming years?

According to Goldman Sachs analysis, token consumption is expected to increase 24 times between 2026 and 2030, reaching around 120 quadrillion tokens a month as more companies adopt AI agents.

How are businesses planning to charge customers for AI-powered services?

Options being explored include raising prices across the board, charging based on completed results, and selling bundles tied to specific tasks or incidents rather than billing purely by token usage.

Related Articles

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Latest Articles