The allure of free AI services, powered by sophisticated Large Language Models (LLMs) like ChatGPT, Claude, and Gemini, presents an undeniable bargain for consumers. Companies such as Microsoft, Google, and Anthropic have poured hundreds of billions of dollars into developing these foundational technologies, making it exceptionally cost-effective for individuals to leverage them for tasks ranging from drafting speeches to planning holidays. However, the economic reality for these tech giants is the imperative to recoup substantial investments, leading them to offer premium, paid-for versions of their AI. These enhanced services provide advanced features crucial for professional applications like coding, complex data analysis, and intricate billing processes.
Beyond the direct offerings of LLM providers, a burgeoning ecosystem of third-party firms is emerging. These companies are building and commercializing services that utilize AI agents, typically underpinned by LLMs, meticulously trained to perform highly specific tasks. Yet, the very act of assigning a price to these AI-driven services is proving to be a formidable challenge. Simon Gooch, from Saviynt, an identity management company integrating agentic AI into its solutions, articulates the difficulty: "Trying to tie someone into a cost model for the next 12 months, two years, three years, it doesn’t make any sense, honestly, because we don’t know." This uncertainty stems directly from the volatile and rapidly evolving economics surrounding "tokens," the fundamental building blocks of LLMs and agentic AI.
At its core, when a user interacts with an LLM – whether asking ChatGPT a question, requesting code from Claude, or automating a process with Gemini – the input prompt is deconstructed into discrete mathematical units known as tokens. These tokens are then processed by the AI model. Similarly, the LLM’s output, whether it’s a textual response, generated code, or a series of commands for automation, is also composed of tokens, which are subsequently translated back into a human-readable or actionable format. The inherent unpredictability of this process is a significant hurdle. Subtle variations in the wording of a prompt can elicit vastly different responses. Furthermore, the same prompt, when submitted multiple times, may not yield identical results, and different LLM models, even when processing the same input, can produce divergent outputs.
The complexity is further amplified in agentic AI systems. Here, businesses orchestrate the collaboration of multiple AI agents, each designed for specific sub-tasks, to collectively make decisions and execute actions. This interconnectedness dramatically increases both the volume of tokens consumed and the overall unpredictability of the output and, consequently, the cost. While the cost per individual token, or the credits used to purchase them, has seen a precipitous decline in recent years, as noted by Goldman Sachs, the sheer quantity of tokens being consumed by both businesses and individual consumers has experienced an explosive surge.
Goldman Sachs forecasts a staggering escalation in token consumption, projecting a 24-fold increase between 2026 and 2030, reaching an astronomical 120 quadrillion tokens per month. This surge is directly attributed to the widespread adoption of AI agents by companies seeking to automate and optimize a growing array of operations. Despite this projected exponential growth, many companies and individuals utilizing AI systems often possess a limited understanding of their actual token expenditure. This lack of awareness typically manifests only when they unexpectedly deplete their allocated tokens or are confronted with unexpectedly high monthly bills.
The financial implications of this unpredictability are already causing concern at the highest levels. Microsoft, a major investor and provider of AI services, has reportedly begun to curb the usage of certain third-party coding tools by its engineers, a move that suggests an internal assessment of AI-related expenditures. Similarly, Uber experienced a dramatic overspend on its AI coding token budget earlier this year, reportedly consuming an entire year’s allocation within a matter of months.
Will Venters, an Associate Professor of Digital Innovation and Information Systems at the London School of Economics, highlights how companies can find themselves caught off guard during the experimental or implementation phases of AI adoption internally. He observes, "People are finding it really hard to manage that cost…it’s a non-deterministic output, so it’s a non-deterministic value." This lack of determinism in AI output directly translates to a lack of determinism in its associated costs, making traditional cost management strategies ineffective.
In response to these challenges, companies are exploring various workarounds. Oliver King-Smith, founder of the engineering software firm smartR AI, suggests that smaller organizations can often operate discreetly by leveraging personal accounts with flat-fee structures. He notes, "This has to end at some point in time, because the big guys are taking a bath on those accounts." King-Smith anticipates that as larger AI platforms face increasing pressure from shareholders to demonstrate profitability, they will inevitably implement stricter controls and pricing adjustments.
King-Smith also emphasizes the importance of strategic decision-making regarding the choice of AI models. Companies need to carefully evaluate which models are most appropriate and cost-effective for their specific needs. Furthermore, Rob Steele, CFO at UK accounting software firm iplicit, stresses the critical need for greater precision in prompt engineering. He draws a relatable analogy: "You wouldn’t send someone in your family out to get the weekly shop without any kind of detailed instructions as to what you expect in that shopping basket, right?" This underscores the necessity of providing clear, detailed instructions to AI models to ensure desired outcomes and minimize wasted computational resources.
The complexity of cost management escalates significantly when AI is integrated into products destined for widespread deployment to thousands of users. Venters points out that the scope of token consumption can rapidly expand beyond core software development to encompass ancillary tasks such as rigorous testing, robust security protocols, and the implementation of critical safety guardrails. "It’s particularly hard when you’re looking at agentic processes," Ventners elaborates. The ease with which more AI agents can be deployed – often with a single click – contrasts sharply with the more deliberate and resource-intensive process of expanding a human workforce, which involves extensive discussions about headcount and hiring.
Venters offers a nuanced perspective, suggesting that while token costs may be unpredictable, the ultimate value derived from token usage with AI might outweigh the perceived expense. He posits, "It’s not quite the same as a calculator. The more you give it, the more expensive it is, but the better the result may be." This highlights the potential for AI to deliver disproportionately high returns on investment, even with fluctuating costs.
However, the fundamental challenge remains for companies to translate these evolving AI costs into sustainable pricing structures for their own customers. Bill Peterson, senior director of product marketing at Sumo Logic, candidly admits, "Nobody’s really figured it out." The software firm is currently in discussions with its corporate clients about how to price its new security services that are built upon agentic AI. "We’re still having some fun conversations about this internally," he remarks wryly.
Peterson outlines potential pricing models, including a general price increase across all offerings, a performance-based pricing structure where customers pay for delivered results, or a model based on "bundles" of specific incidents or services. The inherent instability of this landscape, however, is the constant threat of LLM providers altering their own pricing strategies. "You get into variable pricing, and it’s changing every couple of months," Peterson warns. This volatility is problematic for businesses and their clients, as he concludes, "Customers don’t like that. That’s not how anybody builds a budget." The current pricing predicament for AI services underscores the nascent stage of this transformative technology and the ongoing quest for a stable and predictable economic framework.






