Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

OpenAI Previews Ultrafast Tier: GPT-5.6 Sol Up to 14x Faster

OpenAI has previewed Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14 times faster, reaching as many as 750 output tokens per second. The tier is powered by Cerebras and targets latency-sensitive workloads.

Published

OpenAI previewed Ultrafast on August 13, a new API service tier that runs GPT-5.6 Sol up to 14 times faster than the standard service.

According to the announcement, Ultrafast is powered by Cerebras and delivers up to 750 output tokens per second, a throughput level well above typical API serving.

The tier is aimed at latency-sensitive use cases such as real-time conversation, code completion, and agent applications that need fast round trips.

The partnership with Cerebras signals that OpenAI is exploring dedicated high-speed inference hardware for its flagship models rather than relying on conventional compute scaling alone.

The feature remains a preview, with official pricing and full availability yet to be detailed; developers will be watching the cost and quota policies of the new tier.

The next thing to watch is whether Ultrafast expands to more models and whether high-speed inference becomes a flagship selling point for OpenAI's enterprise API business.

Why it matters

Fast inference is becoming a new battleground for model APIs, and OpenAI's Cerebras-backed Ultrafast preview will test demand from latency-sensitive applications.

OpenAIGPT-5.6CerebrasAPI
Back to realtime news

Nearby Updates

All

08/13, 17:00

OpenAI appoints Dali Rajic as Chief Revenue Officer

OpenAI announced on August 13 that it has appointed Dali Rajic as Chief Revenue Officer to lead the company's global revenue organization and help businesses realize the full value of AI. The executive move underscores OpenAI's deepening focus on enterprise commercialization and revenue-building.

08/13, 19:01

Maverick Payments integrates Findustry AI agent to automate chargeback rebuttals

Payment processor Maverick Payments has integrated Findustry's AI agent to automate chargeback rebuttals, as reported by PYMNTS. The move brings an agent into merchant dispute handling, promising lower operational burden and adding another real-world AI agent deployment in payments.

08/13, 19:29

Claude clears all Hadamard matrices below order 2000, striking another open problem off the math list

Chinese tech outlet QbitAI reports that Anthropic's Claude has cleared all Hadamard matrices below order 2000, removing a long-standing open problem in combinatorics from the to-solve list. The report credits the mathematician's approach as much as the model itself, a fresh sign that large models can contribute to serious mathematical research.

08/13, 19:53

Embodied-data startup SCALEFORCE closes two funding rounds in 40 days to build physical-AI infrastructure

Embodied-intelligence data infrastructure startup Yuanpoint Technology (SCALEFORCE) has completed a new funding round worth tens of millions of yuan, with investors including Hengxu Capital, Kailian Capital and a top domestic embodied-AI industry player — its second round within 40 days. The proceeds will fund the MatrixOS physical-AI operating system, a large-scale data production network, and team expansion.