Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Qwen Office launches Qwen3.8-Flash with standard mode: 100% faster generation, 75% less token usage

Qwen Office launched Qwen3.8-Flash on the evening of August 26, introducing a new standard mode that all users can try immediately. In real-world office testing, the standard mode delivers roughly 100% faster per-task generation and cuts average token consumption by 75%, with 95% of daily tasks expected to be handled in standard mode.

Published
千问办公首发上线Qwen3.8-Flash:生成速度提升100%,Token消耗减少75%
Image source: github.com

On the evening of August 26, Qwen Office (千问办公) launched Qwen3.8-Flash, the model released just earlier that day, alongside a new standard mode. Starting immediately, all users can experience Qwen3.8-Flash through the standard mode.

According to a report by QbitAI (republished with authorization from Qwen), the new model lets users complete tasks with fewer points consumed and faster token throughput. Qwen Office says its future model lineup will have only two tiers — standard and advanced — with 95% of daily tasks handled by the standard mode and only 5% of complex tasks requiring the advanced tier.

The experience gains come from model upgrades combined with agent-side optimization. Built on a new architecture, Qwen3.8-Flash delivers performance surpassing Claude Opus 4.6 with total parameters in the hundreds of billions.

The Qwen large-model team and the Qwen Office team also co-developed an office-specific version of Qwen3.8-Flash, fine-tuned for multi-step planning, tool selection, and context compression, with inference optimization and a custom Harness architecture to further boost throughput.

In real-world office scenario testing, the standard mode improved per-task generation speed by roughly 100% and cut average token consumption by 75%.

The report notes that in real AI applications, high performance usually means high cost and latency, while low cost tends to require sacrificing intelligence — and deep co-optimization between agents and models is breaking this "impossible triangle" of performance, cost, and speed.

As model intelligence density rises and Qwen Office and the models optimize each other, agents are expected to move past token anxiety. The next thing to watch is how the standard/advanced tier pricing plays out and whether the office-specific tuning can extend to more vertical scenarios.

Why it matters

Qwen Office is pushing frontier-model efficiency into everyday office use, and its standard/advanced tiering could reshape how office AI products are priced and experienced.

QwenAlibabaAgent
Back to realtime news

Nearby Updates

All