Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

DeepSeek Routes All V4-Pro API Traffic to V4.1-Flash at Flash Rates

DeepSeek now routes all API traffic aimed at V4-Pro to V4.1-Flash and bills it at Flash rates, according to Pandaily. Requests that need V4-Pro capability are therefore answered by V4.1-Flash, changing the cost picture for callers.

Published
DeepSeek 把 V4-Pro 的 API 流量全部切到 V4.1-Flash,并按 Flash 价格计费
Image source: deepseek.com

DeepSeek has taken an unusual route to a product change. According to Pandaily, the company routes all API traffic aimed at V4-Pro to V4.1-Flash and bills it at Flash rates — a shift that shows up in routing and pricing rather than in a launch event or technical report.

For callers, the practical change is that the path stays the same while the model behind it changes. Requests that require V4-Pro capability are now answered by V4.1-Flash, and the report pairs that substitution with the cheaper billing tier, pointing to a cost reduction on the same endpoints.

Judging by the naming and the Flash rates phrasing, the Flash tier is generally positioned as the lighter, cheaper option. If that reading holds, the net effect is that traffic previously billed at Pro rates is now served by a Flash-class model and charged at Flash prices.

The most notable part is not the model name but how the version change was delivered: not through a blog post or a weight release, but by changing a route. That is fast and efficient, yet it also forces downstream teams to realign benchmarks, cost budgets and model-selection records.

It is worth being clear about the limits of the report. It covers routing and billing only. It does not say whether V4-Pro is being retired, how V4.1-Flash compares on capability, whether there is a schedule, or whether existing integrations need migration. Treating V4-Pro as retired would go beyond what the source says.

From an industry standpoint, quiet switches like this reflect the economics of inference: steering traffic to a cheaper variant inside the same product line is one of the most direct ways to hold a price point. The trade-off is that users have to verify service consistency without a formal announcement.

What to watch next: whether DeepSeek's documentation and pricing pages are updated to match, whether a migration note appears, and what developers measure in latency and quality. Until then, the safe move is to test any V4-Pro integration as if it were a V4.1-Flash integration.

Why it matters

Model version changes are moving from public launches to quiet server-side switches, which means API consumers should treat benchmarking, cost review and migration testing as routine rather than exceptional.

DeepSeekAPI推理服务
Back to AI Daily

Nearby Updates

All

09/14, 11:01

DeepSeek-V4.1-Flash: 552B-Parameter MoE Built for Efficient Inference

A technical breakdown of the DeepSeek-V4.1-Flash model card describes a multimodal mixture-of-experts model with a 552B-parameter backbone that activates only about 8B parameters at a time. The weights ship under an MIT license with a 1M-token context window, but the model still trails the larger V4-Pro family on several reasoning benchmarks.

09/14, 08:30

Zhipu Teases GLM-6.0 in a Financial Filing, Revealing a Fully Self-Trained Approach

Zhipu has slipped the first details of its next flagship, GLM-6.0, into a financial filing rather than a technical blog or paper, according to QbitAI. The filing points to a fully self-trained approach, while the model itself and its accompanying paper have yet to be released.

09/14, 11:18

China's PhysBrain 1.5 tops the global open-source ranking for physical AI

PhysBrain 1.5, a physics-focused AI model from a Chinese team, has reached the top of a global open-source leaderboard, according to QbitAI, which frames the result as clearing the hardest stretch of the physical closed loop. The report puts its spatial intelligence on par with GPT-6 Astra.

09/14, 11:45

CosmosMind and university partners release MetaRSI-v1, an architecture for improving self-improvement

CosmosMind, working with more than ten universities including Stanford, Berkeley, MIT, Tsinghua and Peking University, has released MetaRSI-v1, which it describes as the first architecture to unify Model-RSI, Data-RSI and Harness-RSI. The team also open-sourced its RSI-Harness, reporting an average 10.9-point gain for a 3B-active small model across four benchmarks and 7.3 points for six frontier models on Terminal-Bench 2.1.