Realtime AI News
Cerebras unveils CS-4 inference system, claiming up to 30x faster per-user speed than GPUs
Cerebras has introduced the CS-4, a new AI inference system built from three wafer-scale processors that it says delivers more than 4,400 tokens per second per user and runs up to 30 times faster than GPU solutions. Aimed at agentic workloads, the system begins shipping this quarter.

Cerebras Systems has introduced its latest AI inference system, the CS-4, built from three wafer-scale processors. The company says the system delivers more than 4,400 tokens per second per user on GPT-OSS-120B, an open-weight model with 120 billion parameters, sharpening its case as a potential challenger to Nvidia in the AI infrastructure race.
Cerebras describes the CS-4 as the industry's fastest AI accelerator and the first system built on its next-generation Nexus rack-scale platform architecture. On tokens per second per user, it is almost twice as fast as the CS-3 and up to 30 times faster than GPU solutions given identical prompts.
The design is unusual. Rather than cut a silicon wafer into hundreds of individual chips and connect them on a board, Cerebras keeps the wafer intact and runs it as one processor, with 44GB of memory built onto the wafer itself. Because moving data between a chip and its memory is what slows inference down, the shorter distances matter. The company cites 43.2 petabytes per second of memory bandwidth per wafer, double the previous generation.
The headline metric needs context. Tokens per second per user measures how quickly words arrive for one person waiting on one answer, while throughput measures the total output the system produces for everyone at once, and the two usually pull against each other. GPU clusters group requests into batches to keep hardware busy and cost per token low, but every request in the batch waits its turn; Cerebras is built for the opposite priority. In human terms, 4,400 tokens a second is roughly 3,300 words, arriving faster than anyone can absorb.
Cerebras is betting this speed matters most for agents. "Being 30 times faster doesn't just make a response feel fast. It gives an agentic system room for more than an order of magnitude as much reasoning, verification, or tool use in the same wall-clock time," CTO and co-founder Sean Lie said in the release. Agents run long chains of sequential steps, where each step waits for the previous one and delays stack rather than overlap.
On the spec sheet, Cerebras says compute rises from 125 to 750 petaFLOPS from CS-3 to CS-4, six times the output from three times the silicon. It also claims up to 10 times more throughput per watt, though that comparison is against its own CS-3, not against GPUs. The company moved power conversion roughly 100 times closer to the processor, from about 50 millimetres on conventional GPU boards to about 0.5 millimetres, which it says nearly eliminates board-level power loss.
The "30 times faster than GPU solutions" comparison, however, names no specific GPU. Cerebras does not disclose the vendor, model, or serving configuration of the system it measured against, and the release's own footnote says actual throughput varies by model architecture, context length, precision, and serving configuration.
The release also signals where Cerebras sees itself: its programmable I/O subsystem is pitched as particularly beneficial for disaggregated inference, with AMD Helios and AWS Trainium named as ecosystem partners. In disaggregated setups, one system handles prefill and another handles decode, and Cerebras is positioning itself for the decode side — not as a replacement for GPU clusters but as a component bolted onto one. First shipments begin this quarter.
Why it matters
The CS-4 reframes the inference race around per-user latency for agentic workloads rather than raw throughput; if the claims hold, Cerebras could carve out a niche in a market still dominated by Nvidia.
Nearby Updates
All08/24, 16:53
Inferenz Launches AI Agent That Scores Hospice Eligibility and Prioritizes Admissions
Inferenz has announced general availability of its Hospice Eligibility AI Agent, a new addition to its Caregence agentic AI platform that automates referral eligibility screening and admission prioritization. The agent cuts manual review from 30–90 minutes to about five minutes per referral while keeping a timestamped audit trail for compliance teams.
08/24, 16:07
Alibaba DAMO Academy Launches DAMO LiON, a Liver Cancer AI That Spots 1cm Tumors
Alibaba DAMO Academy, with Shengjing Hospital of China Medical University and other institutions, has launched DAMO LiON, a liver cancer diagnosis AI model that spots tiny lesions on contrast-enhanced CT scans, with results published in Nature Medicine. In a two-month real-world trial, the model caught 15 malignant tumors that radiologists had missed — most around 1 cm — helping patients get timely treatment.
08/24, 15:54
Nvidia in talks to invest in Perplexity at a $30B+ valuation
Nvidia is in talks to invest in AI search startup Perplexity as part of an equity round valuing the company at more than $30 billion, The Information reported on Sunday. The valuation would be up more than 50% from a year ago, as Perplexity's annualized revenue has climbed past $750 million.
08/24, 15:51
Yonyou's H1 2026 Report Signals Enterprise AI Moving into Large-Scale Deployment
Yonyou Network published its 2026 semi-annual report, stating that enterprise AI has entered the stage of large-scale deployment. The signal from one of China's leading enterprise software vendors suggests corporate AI is shifting from pilots to systemic rollout across core business functions.