# Grok 4.6 competes on turns, not IQ — and that is the more interesting claim

> xAI released Grok 4.6 on 12 August 2026 at $2/$6 per million tokens, focused on long-running agents.

*xAI's August release matches GPT-5.6 Sol on a composite index at a third of the headline price. The number that matters is 53 turns against 103.*

By Behzad Hosseini · WireRead
Canonical: https://wireread.com/news/grok-4-6-agentic-efficiency-analysis

The interesting thing about Grok 4.6 is what xAI chose not to claim. There is no assertion of a new frontier, no saturated benchmark, no argument that this is the smartest model available. The launch is explicitly about endurance — models that stay coherent across many steps of a task — and about the cost of getting to the end of one.

## What xAI says it built

The positioning is unusually concrete for a model launch. Rather than leading with a capability ceiling, xAI describes sustained work: researching across many steps, moving through a codebase, turning a rough idea into a working first version.

> builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work
> — [xAI](https://x.ai/news/grok-4-6), 2026-08-12

Two capability notes support that framing. The model is described as especially strong at turning a broad product idea into a working first version — the archetypal agentic task where a model must make a hundred small decisions without supervision. And on extended tasks, xAI reports more self-testing and verification, with the model checking its own work. Self-verification is the quiet unlock for long-horizon agents: a model that notices its own error at step 12 does not compound it through steps 13 to 60.

## The benchmark position, honestly stated

Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index — a nine-benchmark composite — up five points from Grok 4.5 and twenty-three from Grok 4.3 earlier in the cycle. That puts it level with GPT-5.6 Sol and two points behind Claude Opus 5, while holding price flat.

| Model | AA Intelligence Index | Input / output per M |
| --- | --- | --- |
| Claude Opus 5 | 63 | $5 / $25 |
| GPT-5.6 Sol | 61 | $5 / $30 |
| **Grok 4.6** | **61** | **$2 / $6** |
| Grok 4.5 | 56 | $2 / $6 |

Matching a $5/$30 model at $2/$6 is a real commercial result, and it is the headline xAI wanted. On its own ten-row comparison, however, Claude Fable 5 wins the most rows; Grok 4.6 is strongest on knowledge work and legal reasoning, and weakest on terminal use — 26% on Terminal-Bench 3.0, which for an agentic model is an awkward gap, since terminal competence is precisely what long-running engineering agents need.

> matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index
> — [xAI](https://x.ai/news/grok-4-6), 2026-08-12

## Turns are the unit that matters

The most consequential number in the launch is not an index score. On long agentic tasks, Grok 4.6 reportedly finished in roughly 53 turns at $0.84 per task, where Claude Opus 5 took 103 turns for comparable work. If that holds outside the vendor's own harness, it reframes the comparison entirely: a model that reaches the same destination in half the steps is cheaper, faster and — critically for agents — has half as many opportunities to go wrong.

> **Key:** Price per token has become a misleading unit for agentic work. The honest metric is cost per completed task, which multiplies token price by turns taken and failure rate. Every serious lab is now arguing this — and each one draws the boundary where its own model looks best.

One caveat deserves equal weight, because Artificial Analysis flagged it: the upgrade also starts answering later and costs more per measured task than Grok 4.5 in some conditions. A model that thinks longer before its first token can win on total task cost while feeling slower in an interactive session. Which of those you care about depends entirely on whether a human is waiting.

## What comes next, and when

The successor story is messier. Elon Musk described Grok 4.7 as a 2.1-trillion-parameter model that would be 'better than 4.6 in every way, except slightly slower to serve', and it has slipped repeatedly — from 'in 4 weeks' on 24 July, to 'a few weeks', to Musk saying on 11 September that it needed a few more days of reinforcement-learning tuning to fix response-length penalties and task management, noting the model still stops too early on difficult tasks and does not check its own work rigorously enough.

By 13 September he was naming its successor, Grok 4.8, as a 2.5-trillion-parameter model trained on a new C++ stack — while 4.7 remained unreleased. All of that traces to posts on X rather than xAI documentation: no context window, no pricing, no benchmark package, no API model ID. Treat the roadmap as intention, and Grok 4.6 as the model that actually exists. On its merits, it is a credible agentic option at a third of frontier pricing, with a terminal-use gap you should test against your own workload before committing.

## Key takeaways

- xAI released Grok 4.6 on 12 August 2026, one month after Grok 4.5, pitched at long-running agents rather than headline intelligence.
- Pricing is $2 per million input tokens and $6 output, with cached input at $0.50 below 200K prompt tokens and a fast variant at double price.
- It scored 61 on the Artificial Analysis Intelligence Index, up five points on Grok 4.5, matching GPT-5.6 Sol and trailing Claude Opus 5 by two.
- The efficiency figure is the standout: roughly 53 turns at $0.84 per long agentic task, against Claude Opus 5 taking 103 turns for comparable work.
- Terminal work stayed weak at 26% on Terminal-Bench 3.0, and Grok 4.7 has slipped repeatedly since July.

## FAQ

### How much does Grok 4.6 cost?
$2 per million input tokens and $6 per million output, with cached input at $0.50 per million below 200K prompt tokens and a fast variant at twice the standard price.

### Is Grok 4.6 better than GPT-5.6 Sol or Claude Opus 5?
It matches GPT-5.6 Sol at 61 on the Artificial Analysis Intelligence Index and sits two points behind Claude Opus 5 — at roughly a third of their headline token price.

### What is Grok 4.6 best at?
Long-running agentic work, knowledge work and legal reasoning. xAI highlights turning a broad product idea into a working first version, plus more self-verification on extended tasks.

### What is Grok 4.6 weak at?
Terminal use — 26% on Terminal-Bench 3.0 — which matters if your agents do heavy command-line engineering work.

### When is Grok 4.7 coming?
Unclear. It has slipped repeatedly since July and remained unreleased in mid-September 2026, with Musk citing reinforcement-learning tuning and self-checking problems.

## Sources

- [Introducing Grok 4.6](https://x.ai/news/grok-4-6) — xAI, 2026-08-12
- [Grok 4.6: Features, Benchmarks, Pricing & Comparisons](https://www.datacamp.com/blog/grok-4-6) — DataCamp, 2026-08-13
- [xAI's Grok 4.6: Benchmarks, Pricing, and the Full Developer Ecosystem](https://betterstack.com/community/guides/ai/xai-grok-46/) — Better Stack, 2026-08-14
- [xAI delays Grok 4.7 as Musk points to reinforcement-learning and self-checking problems](https://www.datastudios.org/post/xai-grok-4-7-delay-reinforcement-learning-self-checking) — DataStudios, 2026-09-14
- [Space Exploration Technologies Corp — Form S-1 (FY2026)](https://www.sec.gov/Archives/edgar/data/0001181412/000162828026036936/spaceexplorationtechnologi.htm) — U.S. Securities and Exchange Commission, 2026-05-20
- [Anthropic will pay xAI $1.25B per month for compute](https://techcrunch.com/2026/05/20/anthropic-will-pay-xai-1-25-billion-per-month-for-compute/) — TechCrunch, 2026-05-20
