NVIDIA's AVO agent system reaches 100% on ARC-AGI-3, up from a 30% model baseline
NVIDIA published results on August 21, 2026 showing its AVO agent architecture completed the entire ARC-AGI-3 public benchmark set, arguing the achievement demonstrates that agent system design — not raw model capability — is now the binding constraint on long-horizon autonomous performance.
What's new
Per NVIDIA's technical blog, AVO "completed the full 25-environment public set with a 100.00 RHAE score, solving all 183 levels in 6,624 environment actions" on ARC-AGI-3, a benchmark designed to test sustained, long-horizon reasoning rather than single-turn problem solving. NVIDIA reports the underlying model, Claude Opus 5, scores roughly 30% on the same benchmark when used on its own — the full AVO system lifts that to 100% through what NVIDIA describes as persistent memory, supervision, and structured tool-use layered around the model.
NVIDIA also compares AVO's efficiency to VISTA, a prior long-horizon agent system, noting AVO solved the same 183 levels using roughly 12% fewer environment actions (6,624 versus VISTA's 7,542). The company frames AVO as a general-purpose coding agent system built for "sustained autonomous operation across long horizons" rather than a narrow benchmark-specific harness.
Context
ARC-AGI-3 has emerged over the past year as one of the harder public tests of agentic reasoning specifically because it rewards sustained, multi-step planning across many turns rather than single-shot pattern matching, and frontier models alone have historically scored well below saturation on it. NVIDIA's own framing draws a direct line to its recent public argument — echoed in coverage this week — that "the harness, not the AI model, is now the real hero" in agentic AI performance, positioning AVO as evidence for that thesis using a third-party model (Anthropic's Claude Opus 5) as the base.
Why it matters
A jump from a 30% single-model baseline to a 100% full-system score on a benchmark built to resist exactly that kind of improvement is a strong data point for the industry-wide pivot from model-only scaling toward agent-system engineering — memory, supervision, and tool orchestration wrapped around an existing model. That NVIDIA achieved this using a competitor's model (Claude Opus 5) rather than one of its own also underscores that NVIDIA is positioning itself as an infrastructure and systems layer for agentic AI broadly, not just a hardware vendor betting on any single foundation model.
Corroborating sources
- Developer.nvidia
https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/
“AVO completed the full 25-environment public set with a 100.00 RHAE score, solving all 183 levels in 6,624 environment actions”