TechCrunch
Nvidia just showed that the harness, not the AI model, is now the real hero
Nvidia’s research shows that a custom harness—software that manages memory, context, and includes a supervising “CEO-like” component—boosted Claude Opus 5’s score on the ARC-AGI-3 interactive reasoning benchmark from 30% to a perfect 100%. The harness, named Agentic Variation Operators, proved far more decisive than the underlying model for long-horizon tasks, a finding that contrasts with OpenAI’s sub-10% results on the same benchmark.