Skip to main content
WireRead
Back to all news

OpenAI

OpenAI cut Luna by 80% — and told you exactly how it could afford to

A model that helped optimise its own serving stack funded an 80% price cut three weeks after launch. The efficiency story is real; the timing is competitive.

By , Editor-in-Chief · WireReadVerified September 2026

The answer

OpenAI cut GPT-5.6 Luna's price 80% on 30 July, to $0.20 input and $1.20 output per million tokens.

Price cuts in this market usually arrive with no explanation, because the explanation is normally 'a competitor made us'. OpenAI's 30 July cut was unusual in that the company supplied a mechanism, and the mechanism is more interesting than the discount: it says the model helped make itself cheaper to run.

What changed, precisely

GPT-5.6 Luna, the cheapest tier, fell 80% — from $1 and $6 per million input and output tokens to $0.20 and $1.20. Terra, the middle tier, fell 20% to $2 and $12. The flagship, Sol, did not move on price at all; instead it got a faster inference mode delivering 2.5x speed at 2x cost, replacing the previous 1.5x-at-2x option.

Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less, while GPT‑5.6 Terra, our balanced model for everyday work, will cost 20% less.

Source: OpenAI · 30 July 2026

The subscription detail matters more than it looks. OpenAI said the lower prices are also reflected in how usage is counted against paid plans in Codex and ChatGPT Work — so the same Plus or Business seat now stretches roughly five times further on Luna-class work. For teams running agents on a subscription rather than raw API spend, that is the larger practical change.

Tier Before After Change
Luna $1 / $6 $0.20 / $1.20 −80%
Terra $2.50 / $15 $2 / $12 −20%
Sol $5 / $30 $5 / $30 Unchanged (faster mode added)

The self-optimisation loop

OpenAI's stated route to the savings runs across three layers: the models themselves, the inference systems that serve them, and the agentic harness that connects them to tools and context. Inside that, the company describes using GPT-5.6 within its own Codex coding agent to analyse production traffic, rewrite GPU kernels, improve routing heuristics and optimise speculative decoding — reporting roughly 20% lower serving costs from the kernel work and 15% better token-generation efficiency from the decoding changes.

By making every layer more efficient, OpenAI is delivering stronger performance per dollar across more enterprise workloads.

Source: OpenAI · 30 July 2026

This deserves neither breathless framing nor dismissal. It is not recursive self-improvement in the capability sense — no model is making itself smarter here. It is a frontier model doing competent systems-engineering work on a well-specified optimisation problem with a hard, measurable objective, which is exactly the kind of task these models have become good at. The significance is economic: if a lab can point its own model at its serving stack and extract double-digit cost reductions, the cost curve for inference bends from the inside, not just from new silicon.

Why the timing is the story

The GPT-5.6 family launched on 9 July. Twenty-one days later, its cheapest tier lost 80% of its price. Nothing about a three-week-old model's serving economics changes that fast by accident; the surrounding month explains it better. Anthropic launched Claude Opus 5 on 24 July at $5/$25. Google shipped Gemini 3.6 Flash at $1.50/$7.50. And DeepSeek V4 was serving at $0.14/$0.28 while scoring 80.6% on SWE-bench Verified — against which Luna at $1 input looked indefensible for routine work.

That is the shape of the 2026 market: capability differences at the top remain real, but the floor is being set by whoever will serve competent work cheapest, and increasingly that is an open-weight or Chinese provider. Western labs can answer in two ways — cut price, or make the case that price per token is the wrong unit. Both are now happening at once, and the second argument is the one OpenAI leaned on again five weeks later when it priced GPT-6 Astra at $10/$50.

What to watch: whether Terra's more modest cut holds as the mid-tier battleground, whether Sol's margin cushion survives Astra's arrival above it, and whether the self-optimisation loop shows up again as a stated reason for the next cut. If it does, inference cost stops being a function of hardware cycles alone — and starts compounding on model capability.

Frequently asked questions

How much did GPT-5.6 Luna's price drop?
By 80%, from $1/$6 to $0.20/$1.20 per million input/output tokens, effective 30 July 2026. Terra fell 20% to $2/$12.
Did the flagship GPT-5.6 Sol get cheaper?
No. Sol stayed at $5/$30 per million tokens but gained a faster inference mode offering 2.5x speed at 2x cost, up from 1.5x speed at the same premium.
How did OpenAI afford an 80% cut three weeks after launch?
It cites efficiency across models, inference systems and the agentic harness — including using GPT-5.6 in Codex to rewrite production GPU kernels and optimise speculative decoding.
Does the price cut apply to ChatGPT subscriptions?
Indirectly but materially: OpenAI says the lower rates are reflected in how usage counts against paid subscriptions in Codex and ChatGPT Work.
Was this a response to DeepSeek?
OpenAI cites efficiency, not competition. But DeepSeek V4 was serving at $0.14/$0.28, Claude Opus 5 had just launched at $5/$25 and Gemini 3.6 Flash at $1.50/$7.50 — the pressure was unambiguous.

Sources

← All news