OpenAI
OpenAI cut Luna by 80% — and told you exactly how it could afford to
A model that helped optimise its own serving stack funded an 80% price cut three weeks after launch. The efficiency story is real; the timing is competitive.
The answer
OpenAI cut GPT-5.6 Luna's price 80% on 30 July, to $0.20 input and $1.20 output per million tokens.
Price cuts in this market usually arrive with no explanation, because the explanation is normally 'a competitor made us'. OpenAI's 30 July cut was unusual in that the company supplied a mechanism, and the mechanism is more interesting than the discount: it says the model helped make itself cheaper to run.
What changed, precisely
GPT-5.6 Luna, the cheapest tier, fell 80% — from $1 and $6 per million input and output tokens to $0.20 and $1.20. Terra, the middle tier, fell 20% to $2 and $12. The flagship, Sol, did not move on price at all; instead it got a faster inference mode delivering 2.5x speed at 2x cost, replacing the previous 1.5x-at-2x option.
Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less, while GPT‑5.6 Terra, our balanced model for everyday work, will cost 20% less.
The subscription detail matters more than it looks. OpenAI said the lower prices are also reflected in how usage is counted against paid plans in Codex and ChatGPT Work — so the same Plus or Business seat now stretches roughly five times further on Luna-class work. For teams running agents on a subscription rather than raw API spend, that is the larger practical change.
| Tier | Before | After | Change |
|---|---|---|---|
| Luna | $1 / $6 | $0.20 / $1.20 | −80% |
| Terra | $2.50 / $15 | $2 / $12 | −20% |
| Sol | $5 / $30 | $5 / $30 | Unchanged (faster mode added) |
The self-optimisation loop
OpenAI's stated route to the savings runs across three layers: the models themselves, the inference systems that serve them, and the agentic harness that connects them to tools and context. Inside that, the company describes using GPT-5.6 within its own Codex coding agent to analyse production traffic, rewrite GPU kernels, improve routing heuristics and optimise speculative decoding — reporting roughly 20% lower serving costs from the kernel work and 15% better token-generation efficiency from the decoding changes.
By making every layer more efficient, OpenAI is delivering stronger performance per dollar across more enterprise workloads.
This deserves neither breathless framing nor dismissal. It is not recursive self-improvement in the capability sense — no model is making itself smarter here. It is a frontier model doing competent systems-engineering work on a well-specified optimisation problem with a hard, measurable objective, which is exactly the kind of task these models have become good at. The significance is economic: if a lab can point its own model at its serving stack and extract double-digit cost reductions, the cost curve for inference bends from the inside, not just from new silicon.
Why the timing is the story
The GPT-5.6 family launched on 9 July. Twenty-one days later, its cheapest tier lost 80% of its price. Nothing about a three-week-old model's serving economics changes that fast by accident; the surrounding month explains it better. Anthropic launched Claude Opus 5 on 24 July at $5/$25. Google shipped Gemini 3.6 Flash at $1.50/$7.50. And DeepSeek V4 was serving at $0.14/$0.28 while scoring 80.6% on SWE-bench Verified — against which Luna at $1 input looked indefensible for routine work.
That is the shape of the 2026 market: capability differences at the top remain real, but the floor is being set by whoever will serve competent work cheapest, and increasingly that is an open-weight or Chinese provider. Western labs can answer in two ways — cut price, or make the case that price per token is the wrong unit. Both are now happening at once, and the second argument is the one OpenAI leaned on again five weeks later when it priced GPT-6 Astra at $10/$50.
What to watch: whether Terra's more modest cut holds as the mid-tier battleground, whether Sol's margin cushion survives Astra's arrival above it, and whether the self-optimisation loop shows up again as a stated reason for the next cut. If it does, inference cost stops being a function of hardware cycles alone — and starts compounding on model capability.
Frequently asked questions
How much did GPT-5.6 Luna's price drop?
Did the flagship GPT-5.6 Sol get cheaper?
How did OpenAI afford an 80% cut three weeks after launch?
Does the price cut apply to ChatGPT subscriptions?
Was this a response to DeepSeek?
Sources
- Advancing the price-performance frontier with GPT-5.6 — OpenAI, 30 July 2026
- AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost — VentureBeat, 30 July 2026
- OpenAI Cuts GPT-5.6 Luna by 80% Three Weeks After Launch — Enterprise DNA, 31 July 2026
- GPT 5.6 Luna Is 80% Cheaper. The Real Story Is the Price of Thinking — Augmented Mind, 1 August 2026