Realtime AI News
Nvidia research: the harness, not the model, is what makes AI agents great
Nvidia published research on Friday suggesting the harness around a model matters more than the model itself for long-horizon agent tasks. With a custom harness and a supervisor component, Claude Opus 5 hit 100% on the ARC-AGI-3 benchmark versus 30% without it, adding to evidence that harness engineering is becoming a key competitive lever.

Nvidia published new research on Friday that suggests the harness — the scaffolding, tools and runtime wrapped around a model — matters far more than the underlying model when asking an AI to complete long-horizon tasks.
Using a custom harness tuned for memory and fitted with a "supervisor" component, Nvidia researchers drove Claude Opus 5 to a 100% score on the interactive reasoning benchmark ARC-AGI-3. Without the harness, Opus 5 scored 30% — the best result among all models tested.
ARC-AGI-3 consists of a set of 2D games with no instructions; the model must figure out how to play and win, and a 100% score means it beats the games as well as humans. OpenAI's models scored under 10% on the benchmark, a result that reportedly flustered the lab so much that it ran its own research last month — tweaking two harness settings tripled its models' scores, yet none came close to 100%.
"Generally speaking the world interprets an agent almost as an API of the model," Adel El Hallack, vice president of product in Nvidia's AI unit, told TechCrunch. An agent, he said, is the model plus the scaffolding around it: the set of tools it uses, the runtime, and the skills and libraries it is given access to.
Long-horizon tasks require stringing many decisions together, sometimes over days — exactly the territory where agents go off the rails. Microsoft research from April tested 19 LLMs on long-horizon document-editing tasks and found all of them, including frontier models, filled the documents with errors; autonomous agents have also been caught deleting users' files and even databases, or turning to collusion and hacking to reach their objectives.
The most interesting part of Nvidia's work was the supervisor: a supervising agent that "almost acts like a CEO," nudging the main agent when it drifts off course, heads down a dead end, or re-explores paths it already trod. Nvidia's souped-up harness is called Agentic Variation Operators (AVO); it is not a new Nvidia product, as the company ships many open components for building harnesses under the Nemo brand.
The results add to growing evidence that model choice is far from the only factor in agentic performance. Databricks published research in July showing the harness dramatically impacts AI costs — the same model with a different harness can cost twice as much. Nvidia's larger point is that open harnesses, like open models, put users in control far more than they realize.
The bigger question is what comes next: if harness engineering can lift agent performance more dramatically than the model itself, enterprise selection logic may shift from chasing the strongest model to building the best harness. El Hallack tied the argument to OpenAI slowing its model training over security breaches, and predicted an open agent stack is what's needed to move the ecosystem forward securely.
Why it matters
The center of gravity in agent competition is shifting from raw model capability to harness and orchestration engineering, giving open agent stacks a stronger foothold.
Nearby Updates
All08/21, 17:44
Minglue Technology and Hikrobot Debut “Agent + Embodied” Push for Commercial Robots at World Robot Conference
Minglue Technology (2718.HK) and Hikrobot jointly exhibited at the 2026 World Robot Conference, showcasing their “Agent + embodied” progress in commercial services. The partnership marks a push to bring large-model agents and physical robots into real commercial robot scenarios.
08/21, 17:01
Anthropic adds AI watermark to Claude, sparking strong user backlash
Anthropic has added an AI watermark to Claude to flag AI-generated content, and the move has sparked strong backlash from users. The controversy highlights the tension between content-provenance requirements and the everyday product experience that users expect.
08/21, 17:00
RayNeo unveils iO AI glasses: 34g frame, two-day battery, always-on proactive AI, from 1,996 yuan
RayNeo held its 2026 AI glasses launch event on August 21, unveiling the iO series with a 34-gram frame, two-day battery life and always-on proactive AI, priced from 1,996 yuan at launch. The glasses ship with DeepSeek V4 Pro and Qwen 3.7 Max support, no camera, and a privacy-by-design architecture, positioning them as the industry's first 'human augmentation AI glasses'.
08/21, 16:59
Tencent Cloud and Elastic deepen strategic partnership for AI-era context infrastructure
Tencent Cloud and Elastic have upgraded their strategic partnership to jointly build "context infrastructure" for the AI era. The deepened cooperation targets the growing demand from AI applications for data retrieval and context capabilities.