Realtime AI News
Qwen Office launches Qwen3.8-Flash with standard mode: 100% faster generation, 75% less token usage
Qwen Office launched Qwen3.8-Flash on the evening of August 26, introducing a new standard mode that all users can try immediately. In real-world office testing, the standard mode delivers roughly 100% faster per-task generation and cuts average token consumption by 75%, with 95% of daily tasks expected to be handled in standard mode.
On the evening of August 26, Qwen Office (千问办公) launched Qwen3.8-Flash, the model released just earlier that day, alongside a new standard mode. Starting immediately, all users can experience Qwen3.8-Flash through the standard mode.
According to a report by QbitAI (republished with authorization from Qwen), the new model lets users complete tasks with fewer points consumed and faster token throughput. Qwen Office says its future model lineup will have only two tiers — standard and advanced — with 95% of daily tasks handled by the standard mode and only 5% of complex tasks requiring the advanced tier.
The experience gains come from model upgrades combined with agent-side optimization. Built on a new architecture, Qwen3.8-Flash delivers performance surpassing Claude Opus 4.6 with total parameters in the hundreds of billions.
The Qwen large-model team and the Qwen Office team also co-developed an office-specific version of Qwen3.8-Flash, fine-tuned for multi-step planning, tool selection, and context compression, with inference optimization and a custom Harness architecture to further boost throughput.
In real-world office scenario testing, the standard mode improved per-task generation speed by roughly 100% and cut average token consumption by 75%.
The report notes that in real AI applications, high performance usually means high cost and latency, while low cost tends to require sacrificing intelligence — and deep co-optimization between agents and models is breaking this "impossible triangle" of performance, cost, and speed.
As model intelligence density rises and Qwen Office and the models optimize each other, agents are expected to move past token anxiety. The next thing to watch is how the standard/advanced tier pricing plays out and whether the office-specific tuning can extend to more vertical scenarios.
Why it matters
Qwen Office is pushing frontier-model efficiency into everyday office use, and its standard/advanced tiering could reshape how office AI products are priced and experienced.
Nearby Updates
All08/27, 10:46
工业Agent不是“套壳”大模型!西门子百年经验灌进工业AI
工业Agent不是“套壳”大模型!西门子百年经验灌进工业AI. 西门子Xcelerator与普通软件货架最本质的区别。货架解决的是「把产品卖出去」,西门子Xcelerator要解决的是「让产品在真实工业场景中持续生长」。
08/27, 08:24
Viral AI startup Instinct raises $350M at $2.5B valuation
Instinct, the viral AI assistant startup founded just last year, has raised $250 million in a Series B round co-led by Index Ventures and Benchmark, bringing total funding to $350 million at a $2.5 billion valuation. The company, helmed by 23-year-old founder Noah Shinn, remains in private beta while drawing both hype and privacy concerns.
08/27, 07:47
Amazon just tripled its order of Nvidia chips over ‘surging demand’
Amazon just tripled its order of Nvidia chips over ‘surging demand’. Amazon is adding another 2 million Nvidia GPU chips to its data centers over the next two years. But this extended partnerships stretches beyond buying more chips.
08/27, 05:37
Anthropic continues compute gobbling streak in $45 billion deal with Nscale
Anthropic continues compute gobbling streak in $45 billion deal with Nscale. The new deal with the infrastructure provider is the latest example of Anthropic's white hot compute gobbling streak.