Realtime AI News
NVIDIA Vera Rubin NVL72 Claims Up to 30x More Work Per Watt, Setting a New Efficiency Bar for AI Agents
NVIDIA published a blog post claiming its Vera Rubin NVL72 platform delivers up to 30x more work per watt, setting a new efficiency standard for AI agent workloads. The post cites OpenRouter data showing agentic AI workloads consume 15x more tokens than a simple chat request.

NVIDIA published a blog post claiming its Vera Rubin NVL72 platform sets a new efficiency standard for AI agent workloads, with up to 30x more work per watt.
The post cites OpenRouter data showing that agentic AI workloads consume 15x more tokens than a simple chat request, positioning agents as a major driver of inference demand.
Using investment research as an example, the post walks through an agent's workflow: querying financial databases, searching news and regulatory filings, invoking a sub-agent for peer comparisons and valuation modeling, then synthesizing everything into a report.
Such multi-step reasoning, tool calling, and sub-agent collaboration multiply token consumption per task, making cost per token a decisive factor in whether agent applications can scale.
In NVIDIA's framing, the economics of AI factories are defined by delivered output — work per watt, token costs, and utilization — rather than by raw accelerator counts.
Vera Rubin NVL72 is NVIDIA's next-generation rack-scale platform for AI factories, and this efficiency push targets agent workloads, one of the fastest-growing inference scenarios, as the company seeks to defend its position in inference computing.
The question now is whether the 30x figure holds up in third-party benchmarks and real customer deployments, and what it means for cloud purchasing decisions and token pricing.
Why it matters
Agent workloads are becoming a primary driver of inference demand, and NVIDIA's efficiency pitch could intensify competition in the inference market around cost and performance per watt.
Nearby Updates
All08/24, 23:00
Intel Crescent Island GPUs pack up to 32 Xe3P cores, optimized for agentic AI with up to 480GB LPDDR5X
Wccftech reports that Intel's upcoming Crescent Island GPUs pack up to 32 Xe3P cores and are optimized for agentic AI workloads. The lineup uses low-cost LPDDR5X memory reaching up to 480GB of capacity, giving Intel a new angle for large-scale inference deployments.
08/24, 23:00
Nvidia says Groq racks will be online this year following $20 billion purchase
Nvidia says the Groq racks it gained through its roughly $20 billion purchase will be online this year, according to CNBC. The timeline signals that Nvidia is moving its inference infrastructure plans from announcement into deployment.
08/24, 23:00
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
NVIDIA announced that Groq 3 LPX has entered full production and that it is extending its Vera Rubin NVL72 rack-scale system with fast token generation for agentic systems. The company says the next era of AI inference will be defined by how every layer of the AI factory works together.
08/24, 22:45
EverestLabs adds AI agent to its MRF robotics offerings
Waste Dive reports that EverestLabs is adding an AI agent to its robotics offerings for materials recovery facilities (MRFs). Details of the agent's capabilities and deployment timeline have not been disclosed, and the industry is watching AI adoption in recycling automation.