xAI
Grok 4.6 competes on turns, not IQ — and that is the more interesting claim
xAI's August release matches GPT-5.6 Sol on a composite index at a third of the headline price. The number that matters is 53 turns against 103.
The answer
xAI released Grok 4.6 on 12 August 2026 at $2/$6 per million tokens, focused on long-running agents.
The interesting thing about Grok 4.6 is what xAI chose not to claim. There is no assertion of a new frontier, no saturated benchmark, no argument that this is the smartest model available. The launch is explicitly about endurance — models that stay coherent across many steps of a task — and about the cost of getting to the end of one.
What xAI says it built
The positioning is unusually concrete for a model launch. Rather than leading with a capability ceiling, xAI describes sustained work: researching across many steps, moving through a codebase, turning a rough idea into a working first version.
builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work
Two capability notes support that framing. The model is described as especially strong at turning a broad product idea into a working first version — the archetypal agentic task where a model must make a hundred small decisions without supervision. And on extended tasks, xAI reports more self-testing and verification, with the model checking its own work. Self-verification is the quiet unlock for long-horizon agents: a model that notices its own error at step 12 does not compound it through steps 13 to 60.
The benchmark position, honestly stated
Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index — a nine-benchmark composite — up five points from Grok 4.5 and twenty-three from Grok 4.3 earlier in the cycle. That puts it level with GPT-5.6 Sol and two points behind Claude Opus 5, while holding price flat.
| Model | AA Intelligence Index | Input / output per M |
|---|---|---|
| Claude Opus 5 | 63 | $5 / $25 |
| GPT-5.6 Sol | 61 | $5 / $30 |
| Grok 4.6 | 61 | $2 / $6 |
| Grok 4.5 | 56 | $2 / $6 |
Matching a $5/$30 model at $2/$6 is a real commercial result, and it is the headline xAI wanted. On its own ten-row comparison, however, Claude Fable 5 wins the most rows; Grok 4.6 is strongest on knowledge work and legal reasoning, and weakest on terminal use — 26% on Terminal-Bench 3.0, which for an agentic model is an awkward gap, since terminal competence is precisely what long-running engineering agents need.
matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index
Turns are the unit that matters
The most consequential number in the launch is not an index score. On long agentic tasks, Grok 4.6 reportedly finished in roughly 53 turns at $0.84 per task, where Claude Opus 5 took 103 turns for comparable work. If that holds outside the vendor's own harness, it reframes the comparison entirely: a model that reaches the same destination in half the steps is cheaper, faster and — critically for agents — has half as many opportunities to go wrong.
One caveat deserves equal weight, because Artificial Analysis flagged it: the upgrade also starts answering later and costs more per measured task than Grok 4.5 in some conditions. A model that thinks longer before its first token can win on total task cost while feeling slower in an interactive session. Which of those you care about depends entirely on whether a human is waiting.
What comes next, and when
The successor story is messier. Elon Musk described Grok 4.7 as a 2.1-trillion-parameter model that would be 'better than 4.6 in every way, except slightly slower to serve', and it has slipped repeatedly — from 'in 4 weeks' on 24 July, to 'a few weeks', to Musk saying on 11 September that it needed a few more days of reinforcement-learning tuning to fix response-length penalties and task management, noting the model still stops too early on difficult tasks and does not check its own work rigorously enough.
By 13 September he was naming its successor, Grok 4.8, as a 2.5-trillion-parameter model trained on a new C++ stack — while 4.7 remained unreleased. All of that traces to posts on X rather than xAI documentation: no context window, no pricing, no benchmark package, no API model ID. Treat the roadmap as intention, and Grok 4.6 as the model that actually exists. On its merits, it is a credible agentic option at a third of frontier pricing, with a terminal-use gap you should test against your own workload before committing.
Frequently asked questions
How much does Grok 4.6 cost?
Is Grok 4.6 better than GPT-5.6 Sol or Claude Opus 5?
What is Grok 4.6 best at?
What is Grok 4.6 weak at?
When is Grok 4.7 coming?
Sources
- Introducing Grok 4.6 — xAI, 12 August 2026
- Grok 4.6: Features, Benchmarks, Pricing & Comparisons — DataCamp, 13 August 2026
- xAI's Grok 4.6: Benchmarks, Pricing, and the Full Developer Ecosystem — Better Stack, 14 August 2026
- xAI delays Grok 4.7 as Musk points to reinforcement-learning and self-checking problems — DataStudios, 14 September 2026
- Space Exploration Technologies Corp — Form S-1 (FY2026) — U.S. Securities and Exchange Commission, 20 May 2026
- Anthropic will pay xAI $1.25B per month for compute — TechCrunch, 20 May 2026