Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

DeepSeek upgrades V4-Flash into a bargain coding agent model at 28 cents per million output tokens

DeepSeek has retrained its V4-Flash model into a much stronger coding and agent model while keeping API prices near pennies, with output tokens still costing just 28 cents per million. The upgrade scored 82.7 on Terminal-Bench 2.1 and 54.4 on DeepSWE, and Artificial Analysis raised its Intelligence Index rating by 10 points to 50.

Published
DeepSeek升级V4-Flash:输出每百万token仅28美分,编码与Agent能力大幅增强
Image source: deepseek.com

DeepSeek has retrained its V4-Flash model into a far stronger coding and agent model without touching its architecture or its bargain API rates, according to a report from AI newsletter The Neuron on August 2. The publication called it the most important AI upgrade of the weekend, noting that the real story is not a new chatbot trick but the price tag for intelligence reaching a new low.

Rather than scaling the model up, DeepSeek re-trained the existing V4-Flash. The model has 284 billion total parameters but activates only about 13 billion per request, which keeps running costs low. On coding-agent benchmarks, the upgraded model scored 82.7 on Terminal-Bench 2.1 and 54.4 on DeepSWE.

Third-party evaluator Artificial Analysis gave the new V4-Flash a score of 50 on its Intelligence Index, 10 points above the previous Flash model. The Neuron noted the jump is remarkable for a model at this price and size, while cautioning that "Opus-class" claims are benchmark statements, not universal truths about real-world performance.

Pricing is unchanged: $0.14 per million input tokens, $0.28 per million output tokens, and $0.0028 for cached input. That means a million output tokens from V4-Flash cost just 28 cents, which is where the "28-cent agent model" framing comes from. At that rate, large classification jobs, coding loops, and browser-agent retries become feasible without turning every extra attempt into a budget decision.

Developers can use the model as deepseek-v4-flash through DeepSeek's API, self-host it from the Hugging Face repository DeepSeek-V4-Flash-0731, or wait for major U.S. cloud providers to add support. The update also brings support for the Responses API and adaptation for Codex-style coding workflows.

The bigger signal is economic. Most teams do not need the absolute smartest model for every task; they need one that can reliably research, code, classify, or operate tools without making each workflow a luxury purchase. If V4-Flash holds up outside benchmarks, companies could reserve premium models for the hardest judgment calls and route routine agent work to something dramatically cheaper.

The upgrade also pressures OpenAI, Anthropic, and every provider charging a premium for capabilities cheaper competitors are quickly matching. The next model war, The Neuron argues, will be won when buyers start asking why a routine task still requires the expensive option. What to watch next: whether V4-Flash's real-world reliability matches its benchmark scores, and how quickly cloud providers adopt it.

Why it matters

DeepSeek's upgrade turns serious agent workloads into a near-commodity at pennies per task, squeezing premium-priced frontier APIs and making cheap, reliable agent loops the new competitive battleground.

DeepSeekAgentModel Release
Back to realtime news

Nearby Updates

All