Realtime AI News
Chinese Models Lead Weekly Token Volume for 20th Straight Week as DeepSeek V4.1 Flash Hits No. 6
National Business Daily, using the latest OpenRouter data, calculates that global large-model token consumption reached 127 trillion tokens in the week of Sept. 7 to Sept. 13, with Chinese models at 61.17 trillion tokens, leading for a 20th consecutive week. DeepSeek V4.1 Flash, released Sept. 10, climbed to sixth place within three days, while Chinese models took four of the top five slots.
A calculation by the Chinese business daily National Business Daily, based on the latest OpenRouter data, puts global large-model token consumption at 127 trillion tokens in the week of Sept. 7 to Sept. 13, up 10.43 percent week over week. Chinese models accounted for 61.17 trillion tokens, up 7.85 percent, while U.S. models accounted for 21.76 trillion, up 31.56 percent.
Four of the top five models by weekly token volume were Chinese. Tencent's Hunyuan Hy4 preview ranked second with 16.8 trillion tokens, up 15 percent. The report notes that Hy4 preview was officially released and open-sourced on Aug. 28 as a model with 770 billion total parameters and 49 billion active parameters, a context window beyond 1 million tokens, and tuning for agent, coding and productivity workloads.
Zhipu's GLM 5.3 Flash placed third with 11.9 trillion tokens, down 4 percent. DeepSeek V4 Flash 0731, the official release version of V4 Flash, was fourth with 11.6 trillion, down 6 percent, and Xiaomi's MiMo-V2.5 climbed to fifth with 7.77 trillion, up 230 percent.
The week's biggest newcomer was DeepSeek V4.1 Flash. Released on Sept. 10, it reached sixth place within three days of launch with 4.94 trillion tokens. According to the report, V4.1 Flash uses a new Causal-Encoder-Decoder architecture and is a mixture-of-experts model with 552 billion total parameters.
DeepSeek also cut API pricing for V4.1 Flash. Per million tokens, cached input costs 0.02 yuan, uncached input 1 yuan and output 4 yuan during off-peak hours, while peak-hour prices are 0.04 yuan, 2 yuan and 8 yuan. The new prices took effect at 12:00 on Sept. 10, 2026. Compared with V4 Flash at the same hours, cached input is down 60 percent, uncached input about 33.3 percent and output about 11.1 percent.
DeepSeek says V4.1 Flash now surpasses V4 Pro across performance, cost, speed and total task time. Meanwhile MiniMax M3, sixth the previous week, and Zhipu GLM 5.3, ninth, dropped off the list, a sign that turnover inside the ranking has accelerated.
The growth rates tell a subtler story. China's base is far larger but grew 7.85 percent, while U.S. models, at roughly a third of the volume, grew 31.56 percent, suggesting U.S. demand is accelerating again. Twenty weeks at the top does not mean the race has stopped.
Three things to watch: whether V4.1 Flash keeps climbing after its price cut, whether GLM 5.3 Flash and Tencent's Hunyuan Hy4 preview hold their positions, and whether a pricing war spreads to more vendors as agent workloads push token consumption higher.
Why it matters
Token volume is the most direct read on compute demand in the agent era. China's 20-week lead reflects scale built on pricing and open-source strategy, but the 31.56 percent growth on the U.S. side shows the competitive rhythm is still shifting, making DeepSeek's post-price-cut climb a key test of where the pricing war goes next.
Nearby Updates
All09/14, 12:05
OpenAI Pitches AI-Native ChatGPT Ads That Open a Brand Chat, Not a Website
Digiday reports that OpenAI has introduced a new AI-native ad format to select clients, attaching a branded business agent to the ad so that a click opens a chat inside ChatGPT instead of sending users to the advertiser's site. Wayfair is trialing the format, and OpenAI CFO Sarah Friar has described today's response ads as only a "basic starting point."
09/14, 11:01
DeepSeek-V4.1-Flash: 552B-Parameter MoE Built for Efficient Inference
A technical breakdown of the DeepSeek-V4.1-Flash model card describes a multimodal mixture-of-experts model with a 552B-parameter backbone that activates only about 8B parameters at a time. The weights ship under an MIT license with a 1M-token context window, but the model still trails the larger V4-Pro family on several reasoning benchmarks.
09/14, 09:51
DeepSeek Routes All V4-Pro API Traffic to V4.1-Flash at Flash Rates
DeepSeek now routes all API traffic aimed at V4-Pro to V4.1-Flash and bills it at Flash rates, according to Pandaily. Requests that need V4-Pro capability are therefore answered by V4.1-Flash, changing the cost picture for callers.
09/14, 08:30
Zhipu Teases GLM-6.0 in a Financial Filing, Revealing a Fully Self-Trained Approach
Zhipu has slipped the first details of its next flagship, GLM-6.0, into a financial filing rather than a technical blog or paper, according to QbitAI. The filing points to a fully self-trained approach, while the model itself and its accompanying paper have yet to be released.