Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

DeepSeek Routes All V4-Pro API Traffic to V4.1-Flash at Flash Rates

DeepSeek now routes all API traffic aimed at V4-Pro to V4.1-Flash and bills it at Flash rates, according to Pandaily. Requests that need V4-Pro capability are therefore answered by V4.1-Flash, changing the cost picture for callers.

Published
DeepSeek 把 V4-Pro 的 API 流量全部切到 V4.1-Flash,并按 Flash 价格计费
Image source: deepseek.com

DeepSeek has taken an unusual route to a product change. According to Pandaily, the company routes all API traffic aimed at V4-Pro to V4.1-Flash and bills it at Flash rates — a shift that shows up in routing and pricing rather than in a launch event or technical report.

For callers, the practical change is that the path stays the same while the model behind it changes. Requests that require V4-Pro capability are now answered by V4.1-Flash, and the report pairs that substitution with the cheaper billing tier, pointing to a cost reduction on the same endpoints.

Judging by the naming and the Flash rates phrasing, the Flash tier is generally positioned as the lighter, cheaper option. If that reading holds, the net effect is that traffic previously billed at Pro rates is now served by a Flash-class model and charged at Flash prices.

The most notable part is not the model name but how the version change was delivered: not through a blog post or a weight release, but by changing a route. That is fast and efficient, yet it also forces downstream teams to realign benchmarks, cost budgets and model-selection records.

It is worth being clear about the limits of the report. It covers routing and billing only. It does not say whether V4-Pro is being retired, how V4.1-Flash compares on capability, whether there is a schedule, or whether existing integrations need migration. Treating V4-Pro as retired would go beyond what the source says.

From an industry standpoint, quiet switches like this reflect the economics of inference: steering traffic to a cheaper variant inside the same product line is one of the most direct ways to hold a price point. The trade-off is that users have to verify service consistency without a formal announcement.

What to watch next: whether DeepSeek's documentation and pricing pages are updated to match, whether a migration note appears, and what developers measure in latency and quality. Until then, the safe move is to test any V4-Pro integration as if it were a V4.1-Flash integration.

Why it matters

Model version changes are moving from public launches to quiet server-side switches, which means API consumers should treat benchmarking, cost review and migration testing as routine rather than exceptional.

DeepSeekAPI推理服务
Back to realtime news

Nearby Updates

All

09/14, 08:30

Zhipu Teases GLM-6.0 in a Financial Filing, Revealing a Fully Self-Trained Approach

Zhipu has slipped the first details of its next flagship, GLM-6.0, into a financial filing rather than a technical blog or paper, according to QbitAI. The filing points to a fully self-trained approach, while the model itself and its accompanying paper have yet to be released.

09/14, 07:01

The Times: UK ministers avoided AI constraints to protect Britain's national interest

The Times reported on 13 September that British ministers avoided imposing constraints on AI, citing the protection of Britain's national interest as the rationale. The report puts the government's trade-off between safety concerns and industrial competitiveness back in view, making it a signal for where UK rules on advanced AI are heading.

09/14, 05:03

Musk Welcomes Grok's Arrival in Microsoft Copilot as Nadella Touts More Model Choice

Microsoft CEO Satya Nadella announced that xAI's Grok models are being added to Microsoft Copilot, saying the move offers users more model choice, and Elon Musk responded on X to welcome the integration. The rollout initially focuses on Copilot experiences in Word, Excel and PowerPoint, starting with customers in Microsoft's Frontier program.

09/14, 04:10

House Speaker Johnson Rejects an AI 'Moratorium': 'China Will Overlap Us'

House Speaker Mike Johnson told CNN's State of the Union that Congress cannot impose a pause on AI development, saying “China will overlap us” while insisting safety is the industry's own responsibility. President Trump struck the same note the same day, saying the U.S. leads China on AI and that “who wins AI wins,” while allowing that guardrails are possible.