If you have used a free version of ChatGPT or any of its AI rivals, then you are obviously getting a good deal. Firms like Microsoft, Google, and Anthropic have collectively invested hundreds of billions of dollars in developing the sophisticated Large Language Models (LLMs) that power these transformative services. Consequently, accessing ChatGPT, Claude, or Gemini for tasks ranging from speechwriting assistance to holiday planning represents a remarkable bargain for consumers. However, these tech giants are understandably keen to recoup their substantial investments. To achieve this, they offer premium, paid-for versions of their AI, replete with enhanced features tailored for more demanding professional applications like coding and complex financial billing.
Simultaneously, a burgeoning ecosystem of third-party firms is emerging, specializing in building and marketing services that leverage AI agents. These agents, typically built upon foundational LLMs, are meticulously trained to perform specific, often highly specialized, tasks. Yet, a significant hurdle has emerged: the perplexing difficulty in establishing a fair and predictable pricing structure for these advanced AI services. "Trying to tie someone into a cost model for the next 12 months, two years, three years, it doesn’t make any sense, honestly, because we don’t know," admits Simon Gooch, a representative from Saviynt, an identity management company that is actively integrating agentic AI into its offerings.
The root of this pricing quandary lies in the rapidly fluctuating economics surrounding "tokens," which are the fundamental building blocks of LLMs and agentic AI. When a user interacts with an LLM – be it asking ChatGPT for information, prompting Anthropic’s Claude to generate software code, or instructing Gemini to automate a business process – their request is deconstructed into discrete mathematical units known as tokens. These tokens are then processed by the AI model. The AI’s output, whether it’s a textual response, lines of code, or a sequence of commands for process automation, is also generated in the form of tokens, which are subsequently reassembled into human-readable or machine-executable formats.
The inherent unpredictability of this token-based process is a major complicating factor. Subtle nuances in how a user phrases a prompt can lead to vastly different outputs from the AI. Furthermore, the same prompt, when submitted multiple times, may not always yield identical results, and different LLM models are likely to produce distinct answers even for identical queries. The complexity escalates significantly with agentic systems, where multiple AI agents collaborate to make decisions and execute actions. This intricate interdependency further inflates token consumption and introduces a new layer of unpredictability into the overall cost.
While the per-token cost – or the credits used to purchase them – has seen a dramatic decline in recent years, as evidenced by analyses from Goldman Sachs, the sheer volume of tokens consumed by both businesses and individual consumers has experienced an exponential surge. Goldman Sachs forecasts a staggering 24-fold increase in token consumption between 2026 and 2030, projecting a monthly usage of 120 quadrillion tokens as companies increasingly adopt AI agents. However, a significant challenge persists: many companies and individuals using these AI systems possess only a rudimentary understanding of their token consumption, often only becoming aware of their expenditure when they either exhaust their credits or receive a substantial monthly bill.
The financial implications have already begun to surface. Even industry behemoths are feeling the pinch. Microsoft has reportedly curtailed its engineers’ use of certain third-party coding tools due to cost concerns. Similarly, Uber experienced an astonishingly rapid depletion of its annual AI coding token budget within a matter of months earlier this year, highlighting the unforeseen consumption rates. Will Venters, an associate professor of Digital Innovation and Information Systems at the London School of Economics, points out that organizations can be caught off guard during experimentation or internal implementation phases of AI, as employees inadvertently consume vast quantities of tokens. "People are finding it really hard to manage that cost… it’s a non-deterministic output, so it’s a non-deterministic value," he explains, underscoring the difficulty in predicting expenses when the output itself is variable.
In response to these challenges, companies are exploring various strategies to navigate the unpredictable cost landscape. Oliver King-Smith, founder of the engineering software firm smartR AI, suggests that smaller organizations might be able to "fly under the radar" by utilizing personal, flat-fee accounts. While he acknowledges that larger AI vendors likely frown upon this practice, he anticipates that such workarounds will become unsustainable as major players face increasing pressure from shareholders to demonstrate profitability. "This has to end at some point in time, because the big guys are taking a bath on those accounts," he predicts, foreseeing a future where vendors will "start clamping down."
King-Smith also advises companies to adopt a more discerning approach when selecting AI models, suggesting that not all models are created equal in terms of their token efficiency. Rob Steele, CFO at UK accounting software firm iplicit, emphasizes the critical importance of crafting precise prompts. He draws a parallel to everyday life: "You wouldn’t send someone in your family out to get the weekly shop without any kind of detailed instructions as to what you expect in that shopping basket, right?" This meticulous attention to prompt engineering is essential for controlling AI output and, consequently, token consumption.
The complexity of cost management is further amplified when AI is integrated into products destined for widespread deployment to thousands of users, as Venters highlights. The potential for AI costs to balloon becomes a significant concern. Managers may realize that tokens are not only required for core software development but also for ancillary tasks such as rigorous testing, robust security measures, and the implementation of crucial safety guardrails. "It’s particularly hard when you’re looking at agentic processes," Venters notes, explaining that deploying additional AI agents can be as simple as a single click, a stark contrast to the extensive discussions and hiring processes involved in expanding a human workforce.
Venters offers a nuanced perspective, suggesting that while token costs may be unpredictable, the ultimate value derived from AI usage could still be substantial. "It’s not quite the same as a calculator," he posits. "The more you give it, the more expensive it is, but the better the result may be." This implies a potential trade-off between cost and output quality. However, the fundamental challenge for businesses remains the imperative to translate these evolving AI costs into sustainable pricing models for their own customers.
"Nobody’s really figured it out," states Bill Peterson, senior director of product marketing at Sumo Logic. His company is currently previewing new security services built on agentic AI and is actively engaged in discussions with corporate clients regarding pricing strategies. "We’re still having some fun conversations about this internally," he remarks with a touch of wry humor. Potential pricing structures being considered include across-the-board price increases, performance-based pricing (paying for results), or charging for "bundles" of service incidents.
However, Peterson cautions that any pricing structure adopted could be rendered obsolete if and when the major LLM providers alter their own pricing strategies. "You get into variable pricing, and it’s changing every couple of months," he laments. "Customers don’t like that. That’s not how anybody builds a budget." This ongoing flux in underlying costs creates a volatile environment for businesses attempting to establish predictable revenue streams and for customers seeking stable expenditure plans. The nascent stage of AI adoption, coupled with the rapid evolution of the technology and its underlying economics, has created a pricing puzzle that many firms are still struggling to solve.






