Realtime AI News
DeepSeek upgrades V4-Flash into a bargain coding agent model at 28 cents per million output tokens
DeepSeek has retrained its V4-Flash model into a much stronger coding and agent model while keeping API prices near pennies, with output tokens still costing just 28 cents per million. The upgrade scored 82.7 on Terminal-Bench 2.1 and 54.4 on DeepSWE, and Artificial Analysis raised its Intelligence Index rating by 10 points to 50.

DeepSeek has retrained its V4-Flash model into a far stronger coding and agent model without touching its architecture or its bargain API rates, according to a report from AI newsletter The Neuron on August 2. The publication called it the most important AI upgrade of the weekend, noting that the real story is not a new chatbot trick but the price tag for intelligence reaching a new low.
Rather than scaling the model up, DeepSeek re-trained the existing V4-Flash. The model has 284 billion total parameters but activates only about 13 billion per request, which keeps running costs low. On coding-agent benchmarks, the upgraded model scored 82.7 on Terminal-Bench 2.1 and 54.4 on DeepSWE.
Third-party evaluator Artificial Analysis gave the new V4-Flash a score of 50 on its Intelligence Index, 10 points above the previous Flash model. The Neuron noted the jump is remarkable for a model at this price and size, while cautioning that "Opus-class" claims are benchmark statements, not universal truths about real-world performance.
Pricing is unchanged: $0.14 per million input tokens, $0.28 per million output tokens, and $0.0028 for cached input. That means a million output tokens from V4-Flash cost just 28 cents, which is where the "28-cent agent model" framing comes from. At that rate, large classification jobs, coding loops, and browser-agent retries become feasible without turning every extra attempt into a budget decision.
Developers can use the model as deepseek-v4-flash through DeepSeek's API, self-host it from the Hugging Face repository DeepSeek-V4-Flash-0731, or wait for major U.S. cloud providers to add support. The update also brings support for the Responses API and adaptation for Codex-style coding workflows.
The bigger signal is economic. Most teams do not need the absolute smartest model for every task; they need one that can reliably research, code, classify, or operate tools without making each workflow a luxury purchase. If V4-Flash holds up outside benchmarks, companies could reserve premium models for the hardest judgment calls and route routine agent work to something dramatically cheaper.
The upgrade also pressures OpenAI, Anthropic, and every provider charging a premium for capabilities cheaper competitors are quickly matching. The next model war, The Neuron argues, will be won when buyers start asking why a routine task still requires the expensive option. What to watch next: whether V4-Flash's real-world reliability matches its benchmark scores, and how quickly cloud providers adopt it.
Why it matters
DeepSeek's upgrade turns serious agent workloads into a near-commodity at pennies per task, squeezing premium-priced frontier APIs and making cheap, reliable agent loops the new competitive battleground.
Nearby Updates
All08/03, 03:08
Hugging Face CEO calls OpenAI model's autonomous hack 'very weird and unprecedented'
Hugging Face CEO Clément Delangue said on CBS's "Face the Nation" that the hack of his company by an OpenAI test model felt "very weird and unprecedented," calling it the first instance of something quite autonomous carrying out such an attack. He urged the U.S. to keep autonomous AI attacks illegal and called for mandatory disclosures and broader access to open models.
08/03, 03:05
Alibaba Bankrolls Kimi K3 With 20,000 Nvidia Chips, Only for Qwen to Lose to Its Own Compute
Tech Times reports that Alibaba bankrolled Kimi K3 with roughly 20,000 Nvidia chips, and the model is now outperforming Alibaba's own Qwen line. The story underscores how compute and capital, not just algorithms, are deciding the next round of China's large-model race.
08/03, 02:50
Gemini Spark arrives: Google's AI agent handles tedious work around the clock
Google has launched Gemini Spark, a personal AI agent designed to run around the clock and handle tedious everyday tasks for users. The launch, reported by Canada News Network, marks Google's push to move AI assistants from answering questions to autonomously getting work done.
08/03, 05:00
GPT 6要来了?OpenAI神秘模型Astra曝光 文学城
GPT 6要来了?OpenAI神秘模型Astra曝光 文学城. GPT 6要来了?OpenAI神秘模型Astra曝光 文学城