The assumption underneath most AI budgeting has been that token prices fall over time.
In the space of two weeks, one provider cut a flagship tier by 80% and another raised a model’s price by 93%.
The direction is no longer reliable, and the headline rate is increasingly not the rate you pay.
What Moved
OpenAI cut GPT-5.6 Luna from around $1 to $0.20 per million input tokens on July 30, and Luna subsequently became the free ChatGPT default with unlimited chats.
DeepSeek raised V4 Flash from about $0.14 to $0.27 per million on August 14, close to a doubling.
xAI launched Grok 4.6 on August 12 at $2 and $6 per million input and output, matching its predecessor and matching a comparable OpenAI tier on the Artificial Analysis Intelligence Index.
Google reported Gemini crossing one billion monthly active users on August 11.
Those four events describe a market where distribution, not unit cost, is the thing being competed over.
The Long-Context Trap
This is the detail with the most practical consequence, and it is buried in a footnote on most pricing pages.
Grok 4.6 expanded its context window to 500,000 tokens. Pricing doubles to $4 and $12 per million for any request exceeding 200,000 tokens.
The critical part is that the higher rate applies to the entire request, not to the tokens above the threshold.
A 201,000-token request therefore costs roughly twice a 199,000-token request, not marginally more. For anyone processing long documents or running agents that accumulate context, that is a cliff rather than a slope.
Tiered context pricing is becoming common across providers. The question to ask of any long-context model is whether the higher rate applies marginally or retroactively, because the two produce very different bills.
Why Prices Are Rising in Places
Two forces pull in opposite directions.
Efficiency gains reduce the cost of serving a given capability. Optimisus covered one mechanism in the piece on models acing competition maths while missing simple facts, where active parameter counts fell sharply while reasoning scores rose.
Compute costs and capital intensity push the other way. Hyperscaler capital expenditure for 2026 has been projected above $600 billion, and that has to be recovered somewhere.
Free consumer tiers are funded by paid usage elsewhere. A flagship becoming free at the top of the funnel usually means someone further down is paying more.
What This Means for Anyone Building
Three practical consequences.
Model pricing is now an operational variable to monitor rather than a fixed input. A provider that halved its price in July can raise it in September, and DeepSeek just demonstrated that.
Multi-provider architecture is cheaper insurance than it used to be. Several gateways now track hundreds of models across dozens of providers with compatible interfaces, which lowers switching cost substantially.
Context management is a cost control, not just an engineering concern. Trimming a request below a pricing threshold can halve a bill without changing the output.
For crypto and fintech products specifically, where assistants sit over changing market data, the architecture that suits this environment is a small fast model with retrieval rather than a large model with embedded knowledge.
The Regulatory Layer Arriving Alongside
Cost is not the only variable moving. Article 50 of the EU AI Act became enforceable on August 2, requiring chatbots to identify themselves and AI-generated content to be machine-readable.
Optimisus covered what that requires in the piece on support chatbots having to say they are chatbots.
A provider change that alters output marking or disclosure behavior now has compliance consequences as well as cost ones. Those two considerations used to be handled by different teams and increasingly cannot be.
The Honest Summary
Nobody should treat a published token price as durable. The useful discipline is knowing your own cost per task rather than per million tokens, and knowing which threshold your typical request sits under.
Efficiency at the model level is the other half of this, examined in the piece on reasoning scores climbing while parameter counts fall.
The labs are competing on distribution while the unit economics move around underneath. That is a normal phase for infrastructure markets, and it usually ends with fewer providers and firmer prices.
Sources
- AIToolsRecap, AI news August 2026 daily coverage — https://aitoolsrecap.com/Blog/AINewsAugust2026.aspx
- LLM Gateway, new AI model releases August 2026 timeline — https://llmgateway.io/timeline
- Evertune, AI model release tracker — https://www.evertune.ai/resources/ai-model-tracker
- Local AI Zone, latest AI developments August 2026 — https://local-ai-zone.github.io/blog/ai-updates-august-2026.html
- European Commission, safer and more transparent AI — https://commission.europa.eu/news-and-media/news/safer-and-more-transparent-ai-2026-08-02_en
Optimisus covers crypto and technology news for readers who want the detail behind the headline.

