The Token Price Paradox
Vishal Sachar
Co-Founder & CEO of CLRT
Somewhere in your company, two documents disagree. The procurement note says the price of AI fell again this year, as it has every year, by roughly an order of magnitude. The finance report says the AI line on the P&L grew again this year, faster than almost any other cost. Most organisations resolve the contradiction by treating one number as a mistake, or the growth as a phase that discipline will correct. Both numbers are accurate, and the growth is not a phase. The paradox dissolves only when you look underneath the price of the token at the shape of the work the tokens are doing, because the unit of consumption has quietly changed beneath the unit of pricing, and it changed by orders of magnitude.
Start with what actually moved. A chat answer, the workload everyone budgeted for in 2023, is a few hundred tokens: one question in, one answer out. An agent run is a different object entirely. It plans, reasons through intermediate steps, calls tools, reads the results, retries what failed, and checks its own output, and every one of those steps is metered. The largest public window into this shift is the study of 100 trillion tokens of real traffic that a16z, the venture firm, published with OpenRouter: reasoning models carried close to none of the tokens at the start of 2025 and more than half by the end of it, programming grew from roughly 11 percent of tokens to over half, and the average request more than doubled in length in two years. The study's traffic skews toward independent developers and open models, so read it directionally. The direction is not ambiguous.
None of this would matter if prices had held still. They did the opposite. The same firm's LLMflation analysis found the cost of constant-quality inference falling by roughly ten times every year, with GPT-3-class output dropping from $60 per million tokens in late 2021 to six cents three years later, a thousandfold collapse. Stanford's 2025 AI Index measured the same slope another way: the cost of a GPT-3.5-class query fell more than 280-fold in about eighteen months, with declines of nine to 900 times a year depending on the task. Here is the part the market has backwards. This collapse is not the force that will bring your bill down. It is the force that made the bill possible. Work that was uneconomic to automate at $60 per million tokens becomes rational at six cents, so every fall in price recruits new workloads onto the meter. Cheaper tokens do not shrink appetites. They license bigger ones.
The budgeting failure follows from a category error. Two decades of software procurement built its muscle around the licence: a seat costs what it costs, the invoice is boring, and the worst case is paying for seats nobody uses. AI does not price like that. It prices like cloud compute, metered by consumption, elastic in both directions, with spend that follows behaviour rather than headcount. On a meter, two things compound on the same invoice: how many people use the system, and how much each unit of their work consumes. AI is growing along both axes at once, adoption spreading through the organisation while the workloads shift from chat answers to agent runs. A licence budget survives enthusiasm. A metered budget is destroyed by exactly the thing you wanted, which is people using the tools. Companies learned this lesson once, expensively, with cloud bills. AI is rerunning it at a steeper slope.
Uber is the cleanest public proof, precisely because nothing went wrong. Forbes reported in May 2026 that the company had exhausted its 2026 AI budget in four months. The cause was not waste but uptake: the share of engineers using AI coding tools jumped from 32 percent in February to 84 percent in March, and Fortune's account of the internal numbers put spending at $500 to $2,000 per engineer per month. By June, Bloomberg reported, Uber had imposed caps. Read that sequence carefully, because it is the whole argument in miniature. The tools worked, the engineers adopted them faster than any rollout plan assumed, each engineer's consumption ran far past what a per-seat mental model allows for, and the annual budget, an artefact of licence-era thinking, was the only component that failed. The caps are not an embarrassment. They are the first draft of a metering discipline.
So where does that discipline actually live. The instinct is to treat runaway AI spend as a finance problem: set a number, circulate a memo, review monthly. But a memo cannot see where the money moves. The unit of spend is no longer a seat that finance can count. It is a loop that retries, a context that grows, a workflow that calls a frontier model where a smaller one would do, a malformed input that runs all night because nothing told it to stop. Caps imposed from outside the system throttle the people, the one part that was working. The controls that matter are engineering artefacts inside the system: budgets a single run cannot exceed, stop conditions measurable from outside the model's own opinion, spend attributed to workflows rather than teams, so someone can say which of them earns its burn. That last question is not engineering at all. It is judgment, and it is the scarce thing.
Falling token prices are not the force that will shrink your AI bill. They are the force that made it possible.
A deeper dive
The second-order traps are where in-house responses to this stall, and they are worth naming precisely. The obvious fix, a spending cap per team, throttles the valuable half of the equation, adoption, while leaving the wasteful half, the mechanics of each run, completely untouched: the retry loop still retries, the context still bloats, the frontier model still handles work a model a tenth its price could carry. The next fix, a cost dashboard, usually measures the wrong unit, because cost per token is a number nobody can act on while cost per completed unit of work, per resolved ticket, per reviewed contract, per shipped change, is the number that decides whether a workflow deserves to exist, and almost no organisation can produce it. And the subtle trap is asking the system to police itself: an agent asked whether its own run was worth the spend will say yes in a well formatted voice, for the same reason a model grading its own output is not verification. Spend control that works is built from outside the loop, on measurements the loop cannot argue with, at the grain of the workflow rather than the invoice.
This is also why the metering era rewards a different kind of thinking than the licence era did. When software was a licence, the decision was made once, at purchase, and the risk was overpaying. On a meter, the decision is made continuously, by the architecture, and the risk is a workflow that quietly costs more than it returns, forever, at scale, with no single moment where anyone approved it. The companies that get this right will not be the ones with the strictest caps. They will be the ones who know their unit economics per workflow, who route work to the cheapest model that survives verification, whose loops carry budgets and stop conditions the way production code carries tests, and who are willing to conclude that some workflows should not run at all. Every part of that is buildable, but none of it is a memo, and the judgment about which workflows earn their burn is the part that cannot be bought as tooling.
Work with CLRT
If your AI line is growing and nobody can say which workflows earn their burn, the problem is not procurement and a cap will not solve it. It is a question of where AI is pointed and what its unit economics are, workflow by workflow, and that is diagnostic work before it is engineering work. CLRT Ascent was built for exactly this: a structured diagnostic that maps where agentic AI actually pays in your business, at ascent.clrtstudio.com. And when a workflow is worth running, CLRT builds the discipline into the system itself, the run budgets, stop conditions, and attribution that let you scale usage without fearing the meter.

Vishal Sachar
Vishal Sachar is the Co-Founder and CEO of CLRT, where he helps UAE businesses make sense of applied agentic AI and put it to work. He writes on agentic systems, AI governance, and the economics of automation. Reach him at vishal@clrtstudio.com or on LinkedIn.


