Generative AI cost commentary in telecom, and in enterprise IT generally, has mostly stayed anchored to a single number: the price per token, which has been falling steadily as model providers compete and infrastructure efficiency improves. An analysis from Pegasystems published via TM Forum this August makes a case worth taking seriously: falling unit token prices are masking rising total inference spend, and the reason is architectural, not a pricing problem that will solve itself as prices keep dropping.
The Unit Economics Behind Rising AI Spend Despite Falling Prices
Pegasystems’ analysis puts real numbers against the claim, which is unusual in itself — most commentary on enterprise AI cost stays qualitative. The figures cited: approximately $0.016 per 2,400 tokens, scaling to roughly $160,000 per month for a communications service provider running 10 million AI-assisted interactions. The more consequential figure is the breakdown of what those tokens are actually being spent on: roughly 40% of tokens fund retrieval — pulling relevant context, documents, or prior conversation history into the model’s input — rather than the core business logic the interaction is actually trying to accomplish.
That 40% retrieval overhead is the structural explanation for rising cost despite falling per-token prices. As CSPs scale generative AI deployments, the pattern isn’t simply more interactions at the same cost per interaction, it’s agentic loops (an agent taking multiple internal steps to complete a task, each consuming tokens), expanding retrieval scope (pulling more context to improve accuracy, which consumes more tokens per query), and growing context windows (longer conversation or document history carried into each call) all compounding on top of each other. Falling per-token unit prices are a real efficiency gain, but they’re being outpaced by growth in the number of tokens each interaction actually consumes — which is a usage-pattern problem, not a pricing problem, and won’t be solved by waiting for token prices to fall further.
Latency and Auditability as the Under-Discussed Cost Drivers
Pegasystems’ analysis flags two additional pressures that rarely make it into cost discussions focused purely on token spend: latency and auditability. Every additional retrieval step or agentic loop adds latency to the interaction, which has direct customer-experience cost in real-time telecom use cases — customer service, network troubleshooting — independent of the token spend itself. A customer service interaction that takes several additional seconds because the underlying agent is running multiple retrieval steps before responding carries a cost in abandoned interactions and customer satisfaction that doesn’t show up in the token bill at all, but is arguably just as real.
Auditability — being able to reconstruct why an AI system took a particular action or gave a particular answer — becomes harder as agentic loops and retrieval chains grow more complex, which is a governance and compliance cost that compounds alongside the direct compute cost, particularly relevant for telecom operators subject to regulatory scrutiny on automated decision-making. A system with a long, multi-step agentic chain behind a single customer-facing answer is materially harder to audit after the fact than a single-call, single-retrieval interaction, even if both cost roughly the same in tokens.
Decision-Tier Routing as the Recommended Governance Model
The recommendation Pegasystems puts forward is decision-tier routing: pre-invocation governance that routes each interaction to the appropriate processing tier — simple deterministic rules, a smaller and cheaper model, or a frontier large language model — based on the complexity, risk, and business value of that specific interaction, rather than routing every interaction through the most capable, and most expensive, model by default. The logic is straightforward once stated: a routine, low-risk query doesn’t need a frontier model’s full capability and shouldn’t carry its full cost and latency profile; a complex or high-stakes interaction justifies the expense. Most current enterprise AI deployments don’t make this distinction systematically — they either apply a single model tier to everything, or make the tiering decision ad hoc rather than through a governed, pre-invocation routing policy.
What This Means for How Telecom and Industrial AI Buyers Should Model Cost
For any organisation evaluating or already running AI-assisted operations at meaningful scale — whether a telecom operator’s customer service and network operations functions, or an industrial enterprise’s AI-driven maintenance and monitoring systems — the practical takeaway from Pegasystems’ analysis is that cost modelling built purely around per-token or per-API-call pricing will systematically understate real spend as usage scales. A more accurate model accounts for retrieval overhead, agentic loop depth, and context window growth as separate cost drivers layered on top of the base token price, and treats decision-tier routing not as an optimisation to consider later, but as a governance control that should be designed in from the start of any AI deployment expected to scale, with clear ownership of the routing policy itself rather than leaving tiering decisions to whichever engineering team happens to build a given feature.
Questions Worth Putting to an AI Vendor Before Committing to a Pricing Model
The unit economics Pegasystems lays out translate into a specific, practical set of procurement questions that most current vendor conversations skip past in favour of headline per-token or per-call pricing. Buyers evaluating an AI platform or vendor for a scaling deployment should ask what proportion of a typical interaction’s token spend goes to retrieval versus core task completion for their specific use case, whether the platform supports decision-tier routing natively or requires the buyer to build that governance layer themselves, and how the vendor’s pricing model behaves as agentic loop depth grows with more complex tasks over time, rather than assuming a simple, linear relationship between usage volume and cost. A vendor that can answer these questions with real numbers from comparable deployments is offering a materially more trustworthy cost basis for a business case than one whose pricing conversation stays anchored to a headline per-token rate that, on Pegasystems’ own analysis, only tells part of the real cost story.
Explore the full TeckNexus Intelligence Platform — independent, buyer-neutral tools for private network and industrial AI decisions. https://tecknexus.com/intelligence/
















