The 2026 AI Price War Splits in Two: OpenAI Cuts Sol, DeepSeek Raises Prices
OpenAI slashed GPT-5.6 Sol by 20% while DeepSeek hiked V4-Pro up to 11x — the AI pricing war is now moving in opposite directions at once.
For two years, the story of AI pricing only moved in one direction: down. That assumption broke this month. In the span of ten days, OpenAI cut the price of its flagship GPT-5.6 Sol model by more than 20 percent, Google shipped Gemini 3.7 Flash at half the launch price of its predecessor, and DeepSeek — the company that arguably started the price war — raised its own API prices by as much as elevenfold. The frontier model market is no longer racing to the bottom together. It is splitting in two, and which side of the split you land on says a lot about your business model.
OpenAI’s second round of cuts
On August 21, 2026, OpenAI announced it was dropping API and credit pricing for GPT-5.6 Sol, its top-tier frontier model, by over 20 percent for the next three months. The move applies to the API and is rolling out across eligible plans for ChatGPT Work and Codex credits, while Pro, Plus, and Business subscription usage remains unchanged.
The new numbers, confirmed on OpenAI’s developer forum, look like this:
| Model | Input (old → new) | Cached input (old → new) | Output (old → new) |
|---|---|---|---|
| GPT-5.6 Sol | $5.00 → $4.00 | $0.50 → $0.40 | $30.00 → $20.00 |
| GPT-5.6 Terra | $2.50 → $2.00 | $0.25 → $0.20 | $15.00 → $12.00 |
| GPT-5.6 Luna | $1.00 → $0.20 | $0.10 → $0.02 | $6.00 → $1.20 |
Output tokens are the headline: Sol dropped from $30 to $20 per million, a 33 percent cut on the side of the ledger that dominates agent and coding workloads. The promotional pricing is committed through at least November 21, 2026.
This is the second volley in OpenAI’s summer campaign. On July 30 the company cut Luna by 80 percent and Terra by 20 percent, framing it as “advancing the price-performance frontier.” The result was immediate: within roughly two weeks, OpenAI reported crossing one billion users across its models, with GPT-5.6 usage surging after the cuts and the daily text cap being lifted for free users in early August.
The pattern is deliberate. OpenAI is treating price as a product feature, adjusting it in short promotional windows rather than at annual model launches. Each cut is explicitly time-boxed — three months for Sol, three months for the July cuts — which lets the company claim efficiency gains while keeping the option to walk prices back up if demand proves inelastic.
Google matches the blitz
Google is playing the same game from a different angle. Gemini 3.7 Flash, released August 13 and positioned as Google’s “most intelligent workhorse model” for coding and agent work, launched at $0.75 per million input tokens and $3.75 per million output tokens — exactly half of what Gemini 3.6 Flash cost at launch three weeks earlier.
The performance justification is real: Google reports 43.6 percent on FrontierCode 1.1 Main (up from 34.4) and 65.3 percent on DeepSWE v1.1 (up from 49.0), with better instruction-following against design references like screenshots. But the pricing structure carries a catch worth noting: the introductory rate expires December 31, 2026, doubling to $1.50/$7.50 on January 1. Like OpenAI, Google is discounting into a stated window, not cutting forever.
DeepSeek goes the other way
The most striking counter-move came from DeepSeek. On August 13, the Chinese lab moved V4-Pro out of preview into general availability as V4-Pro-0813, posting strong scores — 87.9 on Terminal-Bench 2.1 and 62.7 on DeepSWE, on a reported 1.6-trillion-parameter architecture with a one-million-token context window. Then came the pricing.
DeepSeek introduced peak and off-peak billing, with output tokens during busy hours rising to $3.96 per million, up from a flat $0.87 — an increase of up to 1,100 percent on some token types. Cache-hit pricing jumped from $0.003625 to $0.044 per million. After years of near-free pricing that forced every Western lab to defend their margins, DeepSeek is now charging closer to what serving a frontier-class model actually costs. The company has no cloud business of its own to subsidize token prices, and it is separately reported to be raising close to $8 billion at a roughly $74 billion valuation.
Why the market split
Three structural forces explain the divergence.
Volume players cut; capacity-constrained players raise. OpenAI and Google both own or rent massive inference fleets and monetize usage through subscriptions, enterprise deals, and ecosystem lock-in. Every price cut grows their funnel. DeepSeek, despite enormous demand, faces GPU capacity constraints — Kimi’s Moonshot paused signups last month for exactly this reason — and has no equivalent way to monetize scale beyond raw token revenue. Raising peak-hour prices is demand management disguised as pricing strategy.
Promotional windows are the new list price. None of these cuts are permanent. Sol’s pricing is committed only through late November; Google’s Flash intro rate dies at year-end; OpenAI’s July cuts were similarly bounded. This is airline-style yield management coming to AI: labs are discovering the price elasticity of demand in real time, one window at a time, and reserving the right to re-price upward once usage habits are locked in.
The benchmark gap is too small to justify the price gap. Independent testers note that GPT-5.6 Sol and Anthropic’s Fable 5 are now competitive in day-to-day coding work, with Sol at $4/$20 versus Fable 5 at $10/$50 per million tokens. When capability converges, price becomes the battleground. Anthropic has so far held its prices — a bet that its coding-agent depth justifies the premium — and that bet is now under visible pressure from the discounting side of the market.
What it means for developers
For teams building on frontier APIs, the practical playbook is straightforward. If your workload tolerates model switching, route aggressively: Sol’s promotional window through November 21 and Gemini 3.7 Flash’s intro rate through December 31 are both temporary bargains worth exploiting. Watch the expiry dates — January 2027 is shaping up to be a repricing event across the board, as Google’s rates double and OpenAI’s promos come up for renewal.
If you standardize on one provider, the divergence argues for hedging. The open-weight ecosystem — Alibaba’s Qwen3.8-Max, Meta’s Muse releases, and the pending GLM-5.3 weights — provides a credible negotiating floor. DeepSeek’s price hike is a reminder that “cheap” is a strategy, not a promise, and strategies change when economics demand it.
The comfortable narrative that AI inference would get monotonically cheaper forever was always too simple. What’s emerging instead is a two-tier market: hyperscaler-subsitized volume pricing on one side, and cost-reflective pricing for capacity-constrained specialists on the other. The price war didn’t end this month. It got more interesting.
Sources
- [1] https://openai.com/index/gpt-5-6/
- [2] https://www.reuters.com/technology/openai-cuts-developer-pricing-frontier-gpt-56-sol-model-by-more-than-20-2026-08-21/
- [3] https://community.openai.com/t/20-price-reduction-for-gpt-5-6-sol-api-codex-credits-and-chatgpt-work/1391726
- [4] https://www.cnbc.com/2026/07/30/open-ai-price-cut-gpt.html
- [5] https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/
- [6] https://www.reuters.com/world/china/deepseek-releases-official-v4-pro-model-it-steps-up-expansion-2026-08-13/
- [7] https://unrot.co/blogs/today-top-ai-news-august-23-2026