Realtime AI News
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
NVIDIA announced that Groq 3 LPX has entered full production and that it is extending its Vera Rubin NVL72 rack-scale system with fast token generation for agentic systems. The company says the next era of AI inference will be defined by how every layer of the AI factory works together.

NVIDIA announced on August 24 that it is extending its Vera Rubin NVL72 rack-scale system with fast token generation for agentic systems, and that Groq 3 LPX has entered full production. The company says the next era of AI inference will be defined not by a single breakthrough chip, network, or system, but by how every layer of the AI factory works together.
The announcement positions the Vera Rubin platform as an inference engine for AI agents, which require fast, continuous token generation to sustain multi-step reasoning and tool use. Extending NVL72 with fast token generation is aimed directly at that workload pattern, according to the company.
The second part of the announcement is that Groq 3 LPX has reached full production. Moving from roadmap to shipping scale matters for customers planning data-center deployments around the Vera Rubin generation, because it means the capability is available now rather than promised for later.
NVIDIA framed the update around its broader AI factory strategy, in which infrastructure is designed and built as a complete system rather than a collection of individual accelerators. The same announcement also highlighted the networking and interconnect layers of the AI factory, including NVLink Fusion and Spectrum-X.
Why it matters: agentic AI is shifting the bottleneck in data centers from training throughput to inference latency. Systems that can sustain fast token generation at rack scale are becoming a strategic differentiator for both NVIDIA and the cloud providers building on its platforms.
What to watch next: when Vera Rubin NVL72 systems with the new capabilities become broadly available, and how they perform on real agent workloads compared with the previous generation. For enterprises building agent infrastructure, the production status of Groq 3 LPX is the more immediately actionable signal.
Why it matters
NVIDIA is pushing rack-scale inference hardware toward agent workloads, signaling that fast token generation — not just training capacity — is the next battleground in AI infrastructure.
Nearby Updates
All08/24, 23:00
NVIDIA Vera Rubin NVL72 Claims Up to 30x More Work Per Watt, Setting a New Efficiency Bar for AI Agents
NVIDIA published a blog post claiming its Vera Rubin NVL72 platform delivers up to 30x more work per watt, setting a new efficiency standard for AI agent workloads. The post cites OpenRouter data showing agentic AI workloads consume 15x more tokens than a simple chat request.
08/24, 23:00
Intel Crescent Island GPUs pack up to 32 Xe3P cores, optimized for agentic AI with up to 480GB LPDDR5X
Wccftech reports that Intel's upcoming Crescent Island GPUs pack up to 32 Xe3P cores and are optimized for agentic AI workloads. The lineup uses low-cost LPDDR5X memory reaching up to 480GB of capacity, giving Intel a new angle for large-scale inference deployments.
08/24, 23:00
Nvidia says Groq racks will be online this year following $20 billion purchase
Nvidia says the Groq racks it gained through its roughly $20 billion purchase will be online this year, according to CNBC. The timeline signals that Nvidia is moving its inference infrastructure plans from announcement into deployment.
08/24, 22:45
EverestLabs adds AI agent to its MRF robotics offerings
Waste Dive reports that EverestLabs is adding an AI agent to its robotics offerings for materials recovery facilities (MRFs). Details of the agent's capabilities and deployment timeline have not been disclosed, and the industry is watching AI adoption in recycling automation.