All insights
The Market6 min read

The Price Cut Is the Warning

Vishal Sachar

Co-Founder & CEO of CLRT

Somewhere in your finance system there is a line that improved in August without anyone doing anything. OpenAI cut GPT-5.6 Sol from $5 to $4 per million input tokens and from $30 to $20 on output. Anthropic said the rise it had scheduled for Sonnet 5 would not occur, so the $2 and $10 you were budgeting to lose stayed where they were. Google put Gemini 3.7 Flash on an introductory price that runs to the end of the year. The AI cost line went down, and the temptation is to book it as a saving and move on. It is a saving. It is also the clearest evidence this year that the vendor you built on has stopped selling on capability and started selling on price, and that this happened because buyers like you declined the premium.

41%1
Decline in Ramp's effective price per million tokens index, from a March 2026 peak of $1.15 to $0.68 in early September (card and bill-pay data)
Ramp Economics Lab, 2026
6%2
Share of the tokens businesses purchased from Anthropic in July 2026 that went to Claude Fable 5, the top model, against 11.4 percent of dollars, in Ramp's August AI Index
Ramp Economics Lab, 2026
9.7%3
Fall in median AI spend per employee among the top 1 percent of spending businesses between July and August 2026, from $7,976 to $7,205, in Ramp's September AI Index
Ramp Economics Lab, 2026

On 21 August OpenAI added a line to its GPT-5.6 launch page saying it had dropped the API and credit pricing of Sol by over 20 percent for the next three months; its pricing page now says the promotional rate is available at least through 21 November 2026, a floor date rather than an end date. Anthropic's pricing page says the $2 and $10 per million tokens it had called introductory for Sonnet 5 is now the standard price, and that the previously scheduled increase to $3 and $15 on 1 September will not occur. The same page lists the newer Claude Fable 5.1 at the same $10 and $50 as Fable 5, with cache reads cut from $1 to $0.25 per million. Google's changelog put Gemini 3.7 Flash on an introductory price that expires on 31 December 2026 and made 3.8 Flash generally available on 2 September. Ramp, in its September note, counts a series of price cuts from the two leading labs over a single month.

FIG. 01List prices per million tokens on the vendors' own pages: GPT-5.6 Sol before and after the 21 August promotional cut, and Claude Sonnet 5's scheduled 1 September price against the standard price Anthropic held. Sources: OpenAI and Anthropic pricing pages, September 2026.
01What the buyers did

Ramp is a spend-management company, and its AI Index is built from its own customers' card and bill-pay data, so read it as vendor telemetry rather than a census. Its direction, though, is hard to argue with. The August edition, covering July, found that Claude Fable 5, Anthropic's top generally available model, made up only 6 percent of the tokens businesses bought from Anthropic and 11.4 percent of the dollars. GPT-5.6 Sol, by contrast, was 25 percent of OpenAI's tokens and 23 percent of its spend. Ramp's own explanation was price: Fable cost roughly $10 per million tokens, twice as much as a Sol it described as still highly performant. One month in, in Ramp's words, the buyers who could have paid for it walked past it to the model one rung down.

The September edition, covering August, showed what happened next. Among the top 1 percent of AI spenders, the businesses whose usage the labs most want, median spend per employee fell 9.7 percent, from $7,976 to $7,205 (the September edition restates July's top-tier figure). Ramp's effective price per million tokens index, a reading of what tokens cost rather than of what buyers chose, stood at $0.68, down 41 percent from its March peak of $1.15. Frontier models, which Ramp names as Opus, Fable and Sol, carried 45 percent of token share, down from a 53 percent peak in August. Adoption barely moved: Anthropic's share of US businesses rose 0.34 points to 43.8 percent and OpenAI's rose 0.09 points to 39.8 percent. Put the two editions together and the sequence is plain. Buyers declined the premium in July, the heaviest buyers trimmed in August, and the labs cut. The price cut is not a gift. It is the vendor's reading of your demand curve.

FIG. 02Three Ramp AI Index gauges, each indexed to its earlier reading, with the reported values labelled: effective price per million tokens (March peak to early September), frontier share of tokens (August peak to latest) and top 1 percent spend per employee (July to August). Vendor data. Source: Ramp Economics Lab, 9 September 2026.
02Discipline is not value

It helps to be precise about what a price war does inside a buyer. It lowers the cost of every workflow equally, the ones that earn their tokens and the ones that never did, and it does so without asking which is which. Ramp's concentration figures show how uneven the exposure is. In July, on the August edition's figures, the top 1 percent of businesses spent a median $7,400 per employee on AI, the top 10 percent spent $650, and the median firm spent $11.95. For the median firm the AI bill was never the constraint, and a fifth off $11.95 is not a decision anyone will notice. For the top 1 percent the cut arrives exactly as they are trimming, so the same workflows keep running at a lower rate and the question of which of them earned the money becomes less urgent rather than more. That is the trap. Cheaper tokens subsidise the unmeasured workflow first, because the unmeasured workflow is the one nobody was about to switch off.

FIG. 03Median AI spend per employee in July 2026 by tier of US business, from Ramp's August 2026 AI Index (vendor data; Ramp's September edition restates the top tier and notes that top 1 percent estimates are volatile).

The saving you booked is also less settled than it looks. Sol's rate has a floor date, not an end date, and any budget built on $4 and $20 needs a line for the $5 and $30 that preceded it. Google's Flash pricing is introductory until 31 December. Anthropic's page notes that its newer tokenizer produces roughly 30 percent more tokens for the same text than the one Sonnet 4.6 used, so a per-token comparison with the older model understates the bill for the same work. None of this is deceptive; it is all on the vendors' own pages. It simply means the number you booked is a window, not a price, and the durable saving is not in the rate. It is in stopping the workflows that should not run at any price and doubling the ones that should, and that requires knowing which is which before the window closes.

A price war tells you what the vendor is worried about. It tells you nothing about what your workflows are worth.

A deeper dive

The mechanism is worth walking, because it explains why the cuts landed in a cluster and why they will not be the last. A frontier lab prices its top model on capability and expects the premium to hold while it is the only model that can do certain work. Ramp's July numbers show that premium failing to hold: within Anthropic's own line-up the tokens went to cheaper models, and across vendors they went to Sol at half Fable's price; by September Ramp was describing company-wide defaults that push usage away from the frontier tier toward standard models such as the Sonnet series. Once the share of frontier tokens starts falling, and it fell from a 53 percent peak in August to 45 percent by early September, the lab has two choices. It can wait for a capability gap to reopen, or it can cut. Cutting is faster, and it has the further effect of lowering the effective price across the whole ladder, which is what the 41 percent decline since March records. What looks from outside like generosity is, from inside the lab, a routing decision made by customers, and the labs are chasing the routing. A buyer who understands that will not mistake the next cut for good news either.

The second-order trap sits inside the buyer, not the vendor. A cut converts a value question into a rate question, and finance functions are extremely good at rate questions. Cost per million tokens is a number procurement can benchmark, negotiate and report. Cost per outcome, the number that says whether a workflow returned more than it consumed, requires someone to have named the outcome, measured it, and been willing to hear that a popular pilot returned nothing. When the rate falls, the incentive to ask the harder question falls with it, because the bill got smaller and the person who would have to answer is relieved. Meanwhile the structure of the discount does the vendor's retention work for it: a promotional floor date, an introductory price with a cliff, a cache discount that only pays on workloads shaped a particular way. Each one rewards staying and penalises measuring. Six months from now the rates will have moved again, in either direction, and the firms that treated August as a saving will be renegotiating a bill they still cannot attribute to a single workflow that earned it. The firms that treated it as a warning will be running fewer workflows, on purpose, and will know why.

Work with CLRT

CLRT does not sell tokens, and it has no view on which lab wins the price war. It works on the question the price war leaves untouched: which of your workflows earn their burn at any price, and what it would take to run those to production standard with verification the business can trust. If August's cut made you feel briefly better about a bill you cannot attribute, that feeling is the gap. CLRT Ascent, at ascent.clrtstudio.com, will show you where in your operation the leverage actually sits and what it is worth in dirhams, before the promotional window closes and the rate moves again.

Vishal Sachar

Vishal Sachar is the Co-Founder and CEO of CLRT, where he helps UAE businesses make sense of applied agentic AI and put it to work. He writes on agentic systems, AI governance, and the economics of automation. Reach him at vishal@clrtstudio.com or on LinkedIn.

Start here

Skip the reading. See where your leverage leaks.

Ascent is our free diagnostic. Ten minutes, and you have the one workflow worth building first.