xAI
Grok 4.7 holds its price while three rivals cut theirs — the bill is per task, not per token
xAI's new model beats its predecessor on every benchmark it published, at the same $2/$6 sticker. Artificial Analysis says finishing an average task now takes more than twice the tokens.
The answer
xAI released Grok 4.7 on 21 September 2026 at Grok 4.6's unchanged $2/$6 per million tokens.
xAI released Grok 4.7 on 21 September 2026, and the number that did not move is the interesting one. The model is priced exactly where Grok 4.6 was five weeks ago — $2 per million input tokens, $6 per million output — while the company describes it as a genuinely larger, longer-trained system. Cursor, Grok Build and the Grok API had it live within hours; GitHub began rolling it out across every paid Copilot tier the same day. On the surface, that is a straightforward good-news release: more capability, no price rise. The more useful question, and one xAI's launch page does not answer for you, is what a finished task now costs once every token it actually uses is counted.
What xAI changed under the hood
xAI says Grok 4.7 runs on a new, larger base model than Grok 4.6 and was trained with a longer reinforcement-learning run on a harder mix of tasks, weighted towards problems that take many hours to complete. The company frames the gain as endurance and self-checking rather than raw cleverness — the same framing it used for Grok 4.6 in August.
Grok 4.7 uses a new, larger base model compared to Grok 4.6.
xAI also says it trained the model to natively understand the harness behind Grok Bot, its own agent product, which it credits with better conversational and general-knowledge performance — a sign the lab is now tuning models around its own tooling rather than treating the raw API as the only surface that matters.
It was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete.
The benchmark table, and the index version that changed underneath it
On xAI's own seven-benchmark comparison, Grok 4.7 improves on Grok 4.6 across the board — CursorBench 4.0, DeepSWE v1.1, EEBench, AA Briefcase v1.1, Terminal-Bench 4.0, the Harvey Legal Agent Benchmark and HealthBench Professional. Against Claude Fable 5.1 Max, the picture is split: Fable still leads on CursorBench 4.0 (51.8% to 46.3%), Terminal-Bench 4.0 (57.9% to 37.6%), GDPval (1,735 to 1,695 Elo) and HealthBench Professional, while Grok 4.7 leads Fable on EEBench and the Harvey benchmark.
| Model | Input / output per M tokens | CursorBench 4.0 | Terminal-Bench 4.0 |
|---|---|---|---|
| Claude Fable 5.1 Max | $10 / $50 | 51.8% | 57.9% |
| Grok 4.7 xHigh | $2 / $6 | 46.3% | 37.6% |
| GPT-5.6 Sol Max | $4 / $20 | 41.7% | 37.3% |
| Grok 4.6 High | $2 / $6 | 40.4% | 20.3% |
Independent testing complicates one of those rows. VentureBeat reported that Artificial Analysis, running its own harness rather than xAI's, measured Grok 4.7 at roughly 26% on Terminal-Bench 4.0 at its xHigh setting — well below xAI's own reported 37.6%, and also behind GPT-6 Astra's independently measured 59.6% on the same test. The gap is not necessarily dishonesty; different harnesses, prompts and reasoning-effort settings routinely produce different numbers for the same model. It is, though, a reminder that a vendor's own comparison table is a marketing document with real numbers in it, not a substitute for testing on your own workload.
On Artificial Analysis's Intelligence Index — a ten-evaluation composite now at version 4.3.2 — Grok 4.7 scored 46, two points ahead of Grok 4.6's 44 on the same version. That is a genuine improvement, but it is worth being precise about what it is not: our coverage of Grok 4.6's launch in August quoted it at 61 on the Intelligence Index. That was a different, earlier version of the same index. Artificial Analysis periodically rebuilds the composite's evaluations and reweights them, and scores from different versions are not comparable — a model's number can fall even as its real capability rises, simply because the ruler changed. Always check the version before comparing two Intelligence Index scores.
The number that actually decides the bill
Here is the more consequential finding. Artificial Analysis measured Grok 4.7, at its highest reasoning setting, using roughly 81,000 output tokens to complete an average Intelligence Index task — against about 36,000 for Grok 4.6 at a comparable setting. The rate card did not move. The number of tokens spent reaching an answer did, by well over double.
A model charging less per token can still be more expensive on a finished workload if it needs substantially more reasoning tokens to get there.
Put in dollars, Artificial Analysis's own cost-per-task figures make the point cleanly: Grok 4.7 costs roughly $3.74 to finish an average Index task at its highest setting, and about $2.73 at a lower one — while GPT-5.6 Sol Max, the model xAI's own launch chart benchmarks against at $4/$20 list pricing, comes out at around $1.99 per task on the same measure. The model with the lower sticker price is not, on this evidence, the model with the lower bill. That is the trap in reading a 'price per million tokens' figure as if it were the price of the work.
The timing sharpens the point. A day after Grok 4.7 shipped, Anthropic released Claude Opus 5.5 at $4 input / $20 output per million tokens — 20% below Opus 5 — and OpenAI released GPT-6 Sol at $2/$10 and GPT-6 Luna at $0.10/$0.50, each roughly half the prior generation's rate. Of the three labs that moved inside 48 hours, xAI was the only one that held its price rather than cutting it, betting that a bigger model at the same rate is itself the competitive move. Whether that reads as confidence or as a company with less room to cut depends on how the token-efficiency story lands with buyers running Grok 4.7 at volume.
The practical read for anyone routing production traffic: xAI's rate card is real and unchanged, and the model is a genuine improvement on Grok 4.6 by its own seven-benchmark table. But 'same price, better model' is only half the sentence. The other half is that finishing a job now appears to cost meaningfully more in tokens than it did five weeks ago, at exactly the moment two well-funded rivals cut their own rates. Model the cost of a completed task on your own workload before treating the $2/$6 sticker as the number that matters.
Frequently asked questions
How much does Grok 4.7 cost?
Is Grok 4.7 better than Grok 4.6?
Why does Grok 4.7 score lower than Grok 4.6 did on the Artificial Analysis Intelligence Index?
Does Grok 4.7 cost more to run than Grok 4.6 despite the same price?
How does Grok 4.7 compare on price to Claude Opus 5.5 and GPT-6 Sol?
Sources
- Introducing Grok 4.7 — xAI, 21 September 2026
- Grok 4.7 is now available in GitHub Copilot — GitHub, 21 September 2026
- Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs — Artificial Analysis, 21 September 2026
- Grok 4.7 pairs coding gains with the same affordable pricing — but high token consumption threatens real-world ROI — VentureBeat, 21 September 2026
- Grok 4.7 Brings Big Coding Upgrades to Challenge Claude AI at Unchanged Pricing — Android Headlines, 21 September 2026
- Grok 4.7 pricing: API token costs, tiers, and the Fast tax — eesel AI, 23 September 2026
- xAI launches Grok 4.7 at $2 per million input tokens — tbreak, 22 September 2026