A selection of AI apps on a phone

If you’re used to free versions of ChatGPT, Claude or Gemini, you know the bright side: a quick answer or a bit of code doesn’t cost you a penny. Behind the curtain, however, big tech firms like Microsoft, Google and Anthropic have poured hundreds of billions into building the large language models (LLMs) that power those services. They’ve turned to paid tiers that offer extra features – faster responses, more prompt tokens, or specialised tools for developers and enterprises.

The challenge arrives when third‑party businesses create their own AI agents on top of these LLMs. Each request is sliced into mathematical units called tokens, which the model then processes and converts back into text, code or commands. Because the token count can shift based on wording, length, and even the particular model, predicting the cost per request is far from straightforward.

Simon Gooch of Saviynt – a company embedding agentic AI into its identity‑management services – notes that “trying to tie someone into a cost model for the next 12 to 24 months doesn’t make any sense because we don’t know what the token consumption will be.” When a single prompt can produce a variable number of tokens, even a simple budgeting exercise becomes uncertain.

Goldman Sachs has predicted token consumption could increase 24‑fold between 2026 and 2030, reaching about 120 quadrillion tokens a month as organisations shift to agent‑based solutions. Yet many users have no transparent view of how many tokens they’re actually using until a billing statement arrives.

Some firms are already reacting. Microsoft has reportedly curtailed its internal use of certain third‑party coding tools, while Uber re‑spent a large portion of its AI token budget in just a few months. Others, like SmartR AI’s founder Oliver King‑Smith, are turning to flat‑fee personal accounts to dodge the higher rates, though he warns that “this has to end eventually” as providers tighten policies.

Will Venters of the London School of Economics adds that the unpredictability of token use is especially problematic when an AI‑powered product scales. For instance, training an agent on thousands of users can drive hidden costs for testing, security and guard‑rails that pop up once the service is live.

Venters believes the value gained from AI could outweigh the token cost if managed well, but he stresses the importance of precise prompts to keep usage in check. “If you don’t give detailed instructions, you’re inviting chaos” he says, comparing it to handing someone a shopping basket without a list.

The big question remains: how can businesses incorporate AI’s scalability benefits while maintaining predictable pricing for their customers? Bill Peterson of Sumo Logic, working on AI‑driven security services, is exploring options like charging by results, bundling incidents, or simply raising overall prices – all under pressure to stay competitive as underlying token costs shift with new provider pricing strategies. As the AI economy evolves, the messy interface between token consumption and business revenue will demand smarter economic models and clearer forecasting tools.\u00a0