# Nvidia's AVO agent system solves all ARC-AGI-3 public levels with Claude Opus 5

> Nvidia's AVO system paired with Claude Opus 5 solved all 183 public ARC-AGI-3 levels.

*The same model scored 30% on its own but reached 100% when wrapped in Nvidia's agent harness, researchers said.*

By Behzad Hosseini · WireRead
Canonical: https://wireread.com/news/nvidia-s-avo-agent-system-solves-all-arc-agi-3-public-levels-with-claude-opus-5

Nvidia said on 21 August 2026 that its AVO (Agentic Variation Operators) agent system, running Claude Opus 5, completed all 183 levels across all 25 public environments of the ARC-AGI-3 benchmark. Nvidia said the system worked out what to do with no instructions, rules or stated goals provided.

ARC Prize had previously reported that Claude Opus 5 alone, at high reasoning effort, scored 30.2% on the same public set, according to [The New Stack](https://thenewstack.io/nvidia-avo-arcagi3-benchmark/). Wrapped in AVO, the same model reached 100%.

ARC-AGI-3 is an interactive benchmark of game-like puzzles designed to test whether AI systems can learn new skills the way people do. AVO adds a main agent, persistent memory, tools and a supervisor that intervenes when progress stalls, according to the report.

The write-up was produced by a five-person Nvidia team: Terry Chen, Yeyin Zhu, Zhifan Ye, Jean-Francois Puget and Humphrey Shi.

AVO also showed gains in efficiency, using 6,624 environment actions compared with 7,542 for the rival VISTA harness, a reduction of 12%. The same unchanged agent had previously spent seven days optimising GPU kernels, beating FlashAttention-4 by up to 10.5% on Nvidia's DGX B200 hardware.

Nvidia introduced AVO as a research project in March 2026; it is not a commercial product. The company said configuration differences between the two setups mean the score jump is not a clean measure of AVO's individual contribution.

The public ARC-AGI-3 set is tutorial-like and had already been solved by other harnesses, and no private-set score has been reported. ARC's creator, François Chollet, praised the approach but said, "the private set is the real test of generalisation."

Nvidia has not said when or whether it will report AVO's performance on the private ARC-AGI-3 set.

## Key takeaways

- AVO completed 183 levels across 25 public ARC-AGI-3 environments with no instructions given.
- Claude Opus 5 alone had scored 30.2% on the same public set, according to ARC Prize.
- AVO used 6,624 environment actions, 12% fewer than the rival VISTA harness's 7,542.

## Sources

- [Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia's AVO, it hit 100%](https://thenewstack.io/nvidia-avo-arcagi3-benchmark/) — The New Stack, 2026-08-21
