Realtime AI News
Tencent Hunyuan unveils Hy ASR 3.0 preview, bringing context-aware speech recognition to Yuanbao
Tencent Hunyuan released Hy ASR 3.0 preview on August 4, a speech recognition model that fuses high-precision transcription with Hy3-based semantic understanding, posting WERs of 3.34% for Mandarin, 2.62% for English and 3.12% for Cantonese on open benchmarks. It is now live on Tencent Cloud's API, with the Yuanbao assistant offering the capabilities free of charge.
On August 4, Tencent Hunyuan officially released Hy ASR 3.0 preview, a new-generation speech recognition model built on the language understanding of its latest large language model, Hy3. It fuses high-precision speech recognition with deep semantic understanding, delivering gains across general recognition, context awareness, multi-scenario robustness and dialect coverage, and producing accurate, coherent transcriptions that better match user intent in complex real-world inputs.
On multiple open benchmarks, Hy ASR 3.0 preview keeps word error rates (WER) around 3% across languages — 3.34% for Mandarin, 2.62% for English and 3.12% for Cantonese — leading its competitors. On self-built evaluation sets, it also posts low WERs for general recognition, dialect recognition, context understanding and hard acoustic conditions such as high noise and whispers.
Tencent highlights four capability upgrades: more accurate general recognition for dialects and mixed Chinese-English speech with fewer character errors; better understanding of user intent through contextual homophone correction; easier adaptation to professional scenarios via hotword injection for brand names, people's names and industry terms; and greater stability in complex environments, with targeted optimization for high-noise and whispered speech.
The gains come from architecture, data scaling and post-training working together. Hy ASR 3.0 preview uses an efficiency-focused MoE architecture with the Hy3 base model, plus a self-developed unsupervised speech encoder trained on tens of millions of hours of unlabeled speech to extract high-quality acoustic representations.
The team jointly trained the speech encoder with the LLM on tens of millions of hours of multi-source speech data covering dialects, accents and acoustic environments, refined through a high-quality data pipeline. Multi-stage capability injection, an SFT data system spanning 10 major dialect regions and over 20 second-level sub-regions, and multi-stage reinforcement learning for general transcription accuracy, any-context understanding and long-tail scenarios then sharpened each ability while reducing misrecognition.
Hy ASR 3.0 preview is already available as an API on the Tencent Cloud website, targeting use cases such as intelligent customer service, content understanding and voice search. Tencent's Yuanbao assistant, which co-developed the model, has launched it first — users can press and hold to talk and get dialect recognition, contextual error correction and stable transcription in complex environments, free of charge — with products like WorkBuddy rolling out next.
Why it matters
Hy ASR 3.0 preview pushes speech recognition from transcription toward context-aware understanding, and Tencent's free rollout through Yuanbao puts the capability directly in consumers' hands. Strong dialect coverage and noisy-environment optimization also strengthen Tencent's position in Chinese-language voice AI.
Nearby Updates
All08/04, 17:18
HarmonyOS 7 Smoothes System-Capability Integration: Skills and Agents Become Callable
At the Huawei HDD HarmonyOS Innovation Forum in Xi'an, HarmonyOS 7 showcased a push to package system capabilities for developers: one-tap cross-device transfer, a system-level Xiaoyi agent, and encapsulated Skills and Agents. Developer cases such as Notein and Elephant News report sharply shorter development cycles, with cross-device transfers now under 1.2 seconds.
08/04, 17:22
Mathematicians Reject OpenAI's Claimed Conjecture Breakthrough Within 24 Hours
OpenAI claimed its next-generation model solved ten world-class problems, including overturning the Connes rigidity conjecture. Within a day, mathematician J. L. Nielsen published a rebuttal tracing the 37,000 lines of Lean 4 code and arguing the AI's counterexample fails — the formalization may be internally correct yet irrelevant to the conjecture itself, which remains open.
08/04, 17:36
Alibaba Cloud's Qwen-Image-3.0 Goes Live on Qianwen AI Platform, Tops Domestic Leaderboard
Alibaba Cloud has officially launched Qwen-Image-3.0 on its Qianwen AI platform, making the latest generation of its Qwen image-model line directly available to users and developers. The company also says the model now ranks first among domestic models on an authoritative leaderboard.
08/04, 16:18
DeepSeek's V4 Flash Price Shock Continues: 7-Fen Games and Platforms Racing to Subsidize
DeepSeek's V4 Flash price shock keeps spreading: since the July 31 release of V4-Flash-0731, developers are building games for pennies while platforms like OpenCode, Nous Portal, and Cline pile on subsidies. With frontier-level benchmark scores at US$0.14 per million input tokens, DeepSeek is rewriting developers' default model choices.