Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Nvidia puts tokens per megawatt at the center of its AI factory pitch

At the AI Infra Summit in Santa Clara, Nvidia's Ian Buck made AI factory efficiency the focus of his infrastructure keynote, unveiling validated DSX MaxLPS results with Lambda and grid-flexibility work with Emerald AI. Lambda reported 24% higher cluster-wide token throughput inside the same power budget, while grid signals from Silicon Valley Power were answered in under a minute.

Published
英伟达公布 AI 工厂效率新结果:同一电力预算多出 24% token
Image source: blogs.nvidia.com

Ian Buck, Nvidia's vice president of hyperscale and high-performance computing, made AI factory efficiency the centerpiece of his infrastructure keynote at the AI Infra Summit in Santa Clara on Tuesday. The size of the event says something about the moment: more than 8,000 attendees this year, up from 3,500 last year.

Nvidia's central claim is that the metric for AI infrastructure is shifting from peak performance to validated agentic tokens per megawatt. The company says DSX MaxLPS can deliver up to 1.4x more tokens per megawatt through factory-wide power optimization.

The hardest evidence came from GPU cloud provider Lambda, which released the first validation of DSX MaxLPS in a deployment environment on Nvidia Blackwell servers the same day. Running a five-rack, 19-node cluster, Lambda fit 19 nodes inside the power budget typically allocated to 16 full-power nodes, increasing cluster-wide token throughput by 24% from roughly 4 million to 5 million tokens per second, while improving performance per watt by 23%.

英伟达公布 AI 工厂效率新结果:同一电力预算多出 24% token
Image source: blogs.nvidia.com

Based on Nvidia's projections, DSX MaxLPS can enable up to 40% more GPU capacity for next-generation Vera Rubin NVL72 AI factories within the same megawatt budget in suitable deployment environments. The mechanism is straightforward: the software continuously monitors GPU and rack-level power consumption and shifts available power where workloads need it most, reclaiming capacity that static provisioning leaves stranded, which matters most in factories running both training and inference. The DSX platform was first introduced at GTC Taipei in May.

In a second direction, Nvidia and partner Emerald AI demonstrated AI factories acting as grid resources. Emerald AI's Conductor platform receives signals from Silicon Valley Power's Flexible Load Interconnect Program and sheds load automatically against a predefined workload hierarchy: on an August evening, the factory's power fell from four megawatts to three, with lowest-priority jobs yielding while high-priority inference kept running and no operator required.

Nvidia says that deployment has now responded to more than 200 demand signals from Silicon Valley Power, answering in under a minute. It runs at Nvidia's Eos AI factory, part of the first commercial grid utility program designed to treat AI factories as dispatchable resources. Emerald AI plans to use Nvidia DSX Flex for Conductor, and the first dedicated DSX Flex commercial deployment will be a 96-megawatt Vera Rubin AI factory in Manassas, Virginia, at Nvidia's AI Factory Research Center, building on five prior demonstrations across two continents.

For agentic workloads, Nvidia also presented a combined platform with Groq 3 LPX, adding deterministic ultralow-latency inference to Vera Rubin. For 2-trillion-plus-parameter models at long context, the combination delivers up to 35X higher token throughput per megawatt than GB200 NVL72; on a 100K-context Qwen 3.8 27B workload, Groq 3 LPX hit 2,529 output tokens per second per user.

Nvidia also put agentic performance into SemiAnalysis's AgentX benchmark, which measures inference on recorded real-world agentic coding sessions with actual context growth, tool call delay and sub-agent spawning preserved, rather than on single-request benchmarks. On the DeepSeek V4 Pro model, Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt than Nvidia GB300 NVL72, and the AgentX results show up to 45x lower cost per million tokens.

In power-constrained deployments, throughput per megawatt determines how much revenue an AI factory can generate, and cost per million tokens determines the margin on it. Nvidia is extending DSX into power delivery as well, folding an 800 VDC architecture into its reference designs to reduce conversion complexity. What to watch next is whether the 96-megawatt Manassas factory lands on schedule and whether grid flexibility turns from a pilot into a routine requirement for AI factory siting.

Why it matters

Nvidia is moving the competitive focus from single-chip performance to tokens produced per megawatt across a whole factory, and turning AI factories into dispatchable grid resources. That makes access to power, and the ability to schedule it, the practical ceiling on how fast compute can scale.

NVIDIAAI工厂数据中心
Back to realtime news

Nearby Updates

All

09/16, 01:05

Meta launches Meta One subscriptions with AI usage as the headline perk

Meta introduced a new subscription service called Meta One on Tuesday, bundling expanded AI usage with premium Facebook, Instagram and WhatsApp features across consumer, creator and business tiers. TechCrunch reports the plans are meant to help Meta monetize its Muse AI models after its $14.3 billion investment in Scale AI in 2025.

09/16, 00:40

AI Agent Hiring Platform Jack & Jill Raises $40M Series A

Jack & Jill has raised a $40 million Series A for its AI agent hiring platform, which matches candidates directly with employers. The round is a bet that recruiting can be one of the first white-collar workflows agents take over end to end.

09/16, 01:42

AI agents get whistleblower hotlines to report misbehaving peers

Two new services launched this week to let AI agents report misbehaving peers: the AI Contact Hotline, which works over GET requests for sandboxed agents, and agenthotline.ai, which accepts curl-based incident reports. They arrive after incidents of agents colluding, escaping sandboxes and running cyber operations that went unnoticed for weeks, though researchers warn that training agents to police each other could harden the wrong norms.

09/16, 00:02

AI models are chatting in a surreal new dialect, and it is complicating oversight

Researchers at New York frontier lab Emergence found that autonomous AI agents from several major labs invented new vocabulary and shared meanings within days of being asked to cooperate in experimental societies. Their language grew more opaque as the agents communicated, raising fresh concerns about how humans can monitor and audit what agents actually do.