Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

NVIDIA Vera Rubin NVL72 Claims Up to 30x More Work Per Watt, Setting a New Efficiency Bar for AI Agents

NVIDIA published a blog post claiming its Vera Rubin NVL72 platform delivers up to 30x more work per watt, setting a new efficiency standard for AI agent workloads. The post cites OpenRouter data showing agentic AI workloads consume 15x more tokens than a simple chat request.

Published
NVIDIA Vera Rubin NVL72 宣称每瓦工作量最高提升 30 倍,为 AI 智能体设立效率新标准
Image source: blogs.nvidia.com

NVIDIA published a blog post claiming its Vera Rubin NVL72 platform sets a new efficiency standard for AI agent workloads, with up to 30x more work per watt.

The post cites OpenRouter data showing that agentic AI workloads consume 15x more tokens than a simple chat request, positioning agents as a major driver of inference demand.

Using investment research as an example, the post walks through an agent's workflow: querying financial databases, searching news and regulatory filings, invoking a sub-agent for peer comparisons and valuation modeling, then synthesizing everything into a report.

Such multi-step reasoning, tool calling, and sub-agent collaboration multiply token consumption per task, making cost per token a decisive factor in whether agent applications can scale.

In NVIDIA's framing, the economics of AI factories are defined by delivered output — work per watt, token costs, and utilization — rather than by raw accelerator counts.

Vera Rubin NVL72 is NVIDIA's next-generation rack-scale platform for AI factories, and this efficiency push targets agent workloads, one of the fastest-growing inference scenarios, as the company seeks to defend its position in inference computing.

The question now is whether the 30x figure holds up in third-party benchmarks and real customer deployments, and what it means for cloud purchasing decisions and token pricing.

Why it matters

Agent workloads are becoming a primary driver of inference demand, and NVIDIA's efficiency pitch could intensify competition in the inference market around cost and performance per watt.

NVIDIAVera RubinNVL72AI AgentInference
Back to realtime news

Nearby Updates

All