-+ 0.00%
-+ 0.00%
-+ 0.00%

On the evening of August 26, Smart Spectrum was released. In order to gather extensive and professional feedback from a wide range of users, GLM-5.3-Flash was tested on OpenCode and OpenRouter on an anonymous model OX-Alpha before it was officially released. The Ox-Alpha quickly became the most popular model of the week, setting a new high in dual-platform calls. However, all of these requested traffic is supported by domestic chips with computing power. “In the past week, for the first time, we tried to use large domestic chip clusters to provide services in large-scale traffic. These chips are connected through self-developed high-bandwidth internet networks.” Intellectual spectrum representation. At the cluster level, the company uses a production-grade Encode-Prefill-Decode split architecture to split multi-modal coding, prompt pre-filling, and token-by-token decoding into work pools that can be independently scheduled and independently expandable, thus achieving efficient and reliable services on domestic accelerators. “Compared to the initial baseline on the same hardware, end-to-end service performance has increased by 3 times, and hardware efficiency and single token cost have reached levels comparable to mainstream Nvidia GPUs. This proves that domestic chips can efficiently and economically support the reasoning needs of cutting-edge models in large-scale scenarios.” It's a wise name.

Zhitongcaijing·08/26/2026 23:17:09
Listen to the news
On the evening of August 26, Smart Spectrum was released. In order to gather extensive and professional feedback from a wide range of users, GLM-5.3-Flash was tested on OpenCode and OpenRouter on an anonymous model OX-Alpha before it was officially released. The Ox-Alpha quickly became the most popular model of the week, setting a new high in dual-platform calls. However, all of these requested traffic is supported by domestic chips with computing power. “In the past week, for the first time, we tried to use large domestic chip clusters to provide services in large-scale traffic. These chips are connected through self-developed high-bandwidth internet networks.” Intellectual spectrum representation. At the cluster level, the company uses a production-grade Encode-Prefill-Decode split architecture to split multi-modal coding, prompt pre-filling, and token-by-token decoding into work pools that can be independently scheduled and independently expandable, thus achieving efficient and reliable services on domestic accelerators. “Compared to the initial baseline on the same hardware, end-to-end service performance has increased by 3 times, and hardware efficiency and single token cost have reached levels comparable to mainstream Nvidia GPUs. This proves that domestic chips can efficiently and economically support the reasoning needs of cutting-edge models in large-scale scenarios.” It's a wise name.