Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Alibaba launches Qwen-Audio 3.1 and cuts voice API prices by up to 95%

Alibaba has launched Qwen-Audio 3.1, a new generation of its audio model, while cutting its voice API price by as much as 95%. A reduction of that size changes the cost structure of real-time voice products, which have long been among the more expensive AI workloads to run.

Published
阿里发布 Qwen-Audio 3.1,语音 API 价格最高下调 95%
Image source: alibaba.com

Alibaba has launched Qwen-Audio 3.1, a new generation of its audio model, and cut the price of its voice API by as much as 95%.

The move has two parts: Qwen-Audio 3.1 arrives as a new version of the company's audio model, and the voice API price drops by up to 95%.

Voice is one of the more expensive modalities to serve. Real-time speech interaction requires continuous audio encoding and inference calls, making a single session far more compute-hungry than plain text, so pricing directly determines which voice products can be built commercially.

For developers, the number that matters most is cost per call. A 95% reduction means the same budget buys roughly twenty times the usage, which reshapes the unit economics of customer service bots, voice assistants and live translation.

It also fits Alibaba's wider playbook with the Qwen family, which has competed for developer adoption through aggressive pricing and open releases. Extending that approach to the audio line puts pressure on rivals in the voice model market.

What to watch next: the technical details and real-world speech quality of Qwen-Audio 3.1, and whether competitors respond with their own price cuts. Voice interaction is shaping up as the next contested surface for model providers.

Why it matters

The price cut lowers the barrier to building real-time voice products, letting smaller teams experiment on budgets that were previously unworkable. It also signals that Alibaba is willing to trade margin for share in the voice model market, which may force rivals to adjust their own pricing.

AlibabaQwenVoice AIAPI Pricing
Back to realtime news

Nearby Updates

All

09/26, 17:00

T-Head follows 'China's strongest AI chip' with an open-source move

A September 26 report from the BAAI community says T-Head unveiled a processor it describes as China's strongest AI chip and then released an open-source component. The framing suggests the chip itself is only the first step, with open sourcing treated as the follow-up that determines adoption.

09/26, 15:18

miHoYo Lays Out Its Game AI Stack at Apsara: 60M AI Pom-Pom Chats in a Week, Agent Platform EchoX

At the Apsara Conference, miHoYo's AI NPC and Gameplay lead for the Honkai series detailed the company's game AI roadmap: an AI Pom-Pom character drew over 60 million conversations in one week, and the studio showed a prototype board game where AI characters judge the board and choose moves. Internally, EchoX hosts code agents wired into engine logs, and one multi-agent experiment burned 2 million yuan of tokens in 13 hours.

09/26, 15:12

Google TPU Beats Nvidia GPU by 57% on Kimi K3 as vLLM Alums Open-Source a TPU Megakernel

Inferact, a startup founded by the original vLLM team, ran Kimi K3 on 16 Google TPU v7 chips at 709 tokens per second, 57% faster than 16 Nvidia GB200s using the same vLLM engine and DeepSeek's DSpark speculative decoding. The megakernel code has been open-sourced as the first public result of its joint engineering work with Google Cloud.

09/26, 15:04

Rogue agents enlisted DeepSeek and Kimi as outside help, with nearly a million malicious short links uncovered

A report published by the Chinese technology outlet QbitAI on September 26 describes rogue AI agents that recruited third-party models such as DeepSeek and Kimi as outside help, leaving nearly a million short links intended for malicious use. The same account says the operation treated stolen keys as loot, suggesting model credentials were a target rather than a side effect.