-+ 0.00%
-+ 0.00%
-+ 0.00%

Tencent Hybrid Officially Releases New Generation Speech Recognition Model Hy ASR 3.0 Preview

Zhitongcaijing·08/04/2026 08:57:01
Listen to the news

The Zhitong Finance App learned that on August 4, Tencent Hybrid officially released the Hy ASR 3.0 preview of the next-generation voice recognition model. Currently, Hy ASR 3.0 Preview has been launched on Tencent Cloud's official website to provide API services, which can be widely used in scenarios such as intelligent customer service, content understanding, and voice search. Yuanbao is deeply involved in joint research and has completed the initial launch. Users can experience improved capabilities such as dialect recognition, intelligent contextual error correction, and stable transcription in complex environments when opening Yuanbao to speak. Voice input is more accurate, stable, and free; products such as WorkBuddy are also being added one after another.

Based on the language comprehension capabilities of Hy3, the latest-generation big language model, Hy ASR 3.0 preview combines high-precision speech recognition with deep semantic understanding, and achieves comprehensive improvements in core dimensions such as general recognition, context perception, multi-scene robustness, and dialect coverage. It can provide accurate, coherent, and closer to user intent in more complex real input, and evolved from “word-by-word transcription, single-point optimization” to “understanding context, compatible scenarios, and straight out with one click”.

Hy ASR 3.0 preview led the overall performance in multiple open source review sets and self-built review sets. In the open source review, Hy ASR 3.0 preview controlled the WER (Word Error Rate) of multiple languages at around 3%, including Mandarin Chinese WER 3.34%, English WER 2.62%, and Cantonese WER 3.12%.

From the perspective of user use, Hy ASR 3.0 preview focuses on improving four types of capabilities —

More accurate universal recognition: Improves recognition accuracy in scenarios such as common speech, dialects, and mixed Chinese and English speech, and further reduces the accumulation of errors in typos, omissions, and long audio.

Can better understand user intent: combine context with Context to accurately capture user context, intelligently correct homophones, and eliminate semantic ambiguity.

Easier to adapt to professional scenarios: Supports enhanced hot word injection to help models quickly identify brand product names, names and industry terms, and reduce business access and ongoing maintenance costs.

More stable in complex environments: Specially optimized for various acoustic scenarios such as high noise, whispering, and quiet speech, stable performance is maintained even under complex conditions.