Zhitong Finance App News, Yun Zhisheng (09678) announced that recently, the company has fully completed the U2-ASR and U2-TTS capability upgrades, continued to strengthen multi-modal large model capabilities, and achieved extensive expansion from “100 dialects” to “global multilingual” in the field of large voice models, further consolidating global voice interaction infrastructure capabilities, and providing efficient and low-threshold technical support for scenarios such as cross-border overseas, physical intelligence and agent interaction.
In this upgrade, U2-ASR added 13 new international language recognition capabilities, covering key overseas markets such as Europe, Southeast Asia, the Middle East and Latin America; U2-TTS added speech synthesis capabilities for 8 Southeast Asian languages. So far, the U2 voice model has supported more than 100 Chinese dialects and more than 15 international languages. Enterprises only need to connect to one model and one set of interfaces to process audio content in different languages, greatly lowering the threshold for the development, deployment and maintenance of multilingual voice services.
Under a unified evaluation scale, U2-ASR has excellent comparative performance with the industry benchmark model: the average typographical error rate (CER) for 113 languages is only 6.58%. In real business scenarios where language tags are not imported, it still maintains a high accuracy rate with automatic language recognition and closed collection routing capabilities, effectively avoiding language misjudgment. Meanwhile, in the ChinaVoices Challenge 2026 Chinese Multi-dialect Speech Recognition Challenge, U2-ASR received leading results in both the first place for the limited data plan and the second place for the open data plan with CER 9.235%, further verifying the leadership of its speech recognition technology. U2-TTS leads the mainstream model in comprehensive performance in subjective evaluation of comprehensibility and naturalness, and uses a streaming neural network acoustics model to achieve real-time output of 24kHz high-fidelity voice by chunk. The first packet delay is significantly reduced to meet the requirements of time-efficient interaction such as real-time voice conversations.
This upgrade enables U2-ASR and U2-TTS to seamlessly collaborate to build a complete multilingual voice interaction link for enterprises (Agents). Voice technology is gradually becoming the core interaction infrastructure for global businesses. Currently, the upgraded U2-ASR and U2-TTS have been fully launched on the company's TokenHub large model service platform and opened standard APIs. The company will continue to implement the concept of “intelligence for good”, break down language barriers with the power of technology, promote inclusive sharing of AI technology, and enable all countries and regions at different stages of development to enjoy the technological dividends of the AGI era equally.