The Zhitong Finance App learned that the AI chip company Cerebras Systems (CBRS.US) officially launched the next-generation rack-level AI system Cerebras CS-4 at the SUPERNOVA press conference held in San Francisco, claiming that it is “the fastest AI accelerator in the industry.” The system targets the AI inference circuit and directly targets Nvidia (NVDA.US) GPU server products. Shipments of the CS-4 are expected to begin this quarter (third quarter 2026), and samples are currently being tested by a small number of customers.
Tech specs: Three wafer-level processors deliver 750 pFlops of computing power
CS-4 is the first product of Cerebras' next-generation Nexus rack-scale platform architecture, driven by three new Wafer Scale Engine 3 Turbo (WSE-3T) wafer-level processors. The WSE-3T is the largest AI processor to date. Each chip integrates 4 trillion transistors and 900,000 AI-optimized cores, covers an area of 46,225 square millimeters of silicon, and has a built-in 44GB SRAM.
The CS-4 system provides 750 pflops of AI computing power, 129.6 PB of memory bandwidth per second, and 7.2 Tbps of I/O throughput per second. Compared to the previous generation CS-3, the CS-4 is twice as fast and has a 10-fold increase in throughput per watt, significantly improving the economic efficiency of the data center. Compared with the previous generation WSE-3, the WSE-3 doubles the AI computing power of a single wafer from 125 PFLOPS to 250 PFLOPS, and the memory bandwidth from 21.6 Pb/s to 43.2 Pb/s. The power consumption of a single chip also jumped from 15 kilowatts to around 33 kilowatts.
In terms of architecture design, CS-4 uses a modular Nexus platform architecture that separates computing, power, and I/O as independent components. The power conversion module was shortened from about 50 mm from the processor on the traditional GPU board to about 0.5 mm, almost eliminating board-level power loss. The pluggable “backpack” design integrates power conversion, liquid cooling, and control electronics, reducing deployment time from days to hours.
Performance breakthrough: more than 4400 tokens per second per user in GPT-OSS-120B tests
Cerebras claims that the CS-4 has set a new industry benchmark in the speed of reasoning. In direct comparison tests on the GPT-OSS-120B model, CS-4 can generate more than 4,400 tokens per second per user, up to 30 times that of the GPU solution. The system can support models with more than 50 trillion parameters.
Sean Lie, chief technology officer and co-founder of Cerebras, said, “30x faster doesn't just make the response feel smoother; it allows intelligent systems to have more than an order of magnitude of space to reason, verify, or invoke tools in the same actual time.” Cerebras CEO Andrew Feldman said, “Historically, quick reasoning meant using smaller, less capable models. The CS-4 achieved industry-leading speed on the largest cutting edge model, fundamentally changing this paradigm”.
In terms of energy efficiency, the CS-4's throughput per watt is up to 10 times higher than the CS-3, significantly improving the economic efficiency of the data center. Cerebras CEO Andrew Feldman pointed out that since fast tokens are more valuable than slow tokens, CS-4 can provide higher quality tokens and more total tokens within a given power budget, “fundamentally changing the paradigm.”
CS-4 also uses a design that reduces the number of components by 50%, making customer deployment easier and faster. In terms of system architecture, the new Nexus platform separates computing, power supply, and I/O into modular components — the power supply conversion was reduced from about 50 mm from the traditional GPU motherboard to about 0.5 mm, drastically reducing board-level power consumption and supporting higher operating frequencies. The pluggable “backpack” design integrates power conversion, liquid cooling, and control electronics in a separate unit, reducing deployment time from days to hours.
Notably, Cerebras' wafer-level systems use static random access memory (SRAM) rather than industry-common DRAM. SRAM is much faster than DRAM, but is larger and more expensive — Cerebras dinner-sized wafers provide the physical space needed to utilize SRAM. At the same time, the single-chip design makes the data transmission distance much shorter than Nvidia or AMD's multi-chip systems.
Ecological Cooperation: Teaming up with AMD Helios to improve efficiency by 5 times
CS-4's programmable I/O subsystem supports the Ethernet-based RoCE v2 RDMA standard and can be integrated with heterogeneous systems. Cerebras has partnered with AMD (AMD.US) on this architecture, combining AMD's Helios rack-level solution with WSE, which can generate 5 times more tokens per watt than a pure Cerebras configuration. AWS Trainium is also listed as a partner.
Dylan Patel, founder and CEO of SemiAnalysis, commented that CS-4's improvements in deployability and network capabilities enable it to expand to larger models and build a “large-scale token factory.”
Market background: After listing, the stock price fell 35% from a high point, and CS-4 became a key turning point
On May 14, 2026, Cerebras was listed on NASDAQ at a price of 185 US dollars per share. The opening price on the first day was 350 US dollars, and the closing price was 311.07 US dollars, an increase of 68%. However, since then, the stock price has continued to fall. It was reported at around $220 on Tuesday, down about 35% from its high point.
The company's second-quarter earnings report was mixed: core revenue of US$209.9 million, up 103% year over year; however, the loss per share was US$2.98, compared to a profit of US$1.91 for the same period last year. The company's guidance for a tripling of core revenue growth in 2027 and a Q3 outlook that exceeded expectations did not fully impress investors. However, the company's remaining performance obligations amount to US$25.4 billion, providing strong visibility into future revenue.
Industry significance: The AI reasoning market pattern is facing reshaping
The release of CS-4 comes at a critical point where demand for AI inference is exploding. As large models move from the training stage to large-scale deployment, inference efficiency is becoming the core competitive dimension of AI infrastructure. Cerebras's claimed 30x speed advantage, if validated in independent benchmarks, could change the purchasing decisions of hyperscale data center operators.
The release of Cerebras CS-4 indicates that AI chip competition is expanding at an accelerated pace from model training to the inference market. Whether the system can truly shake up Nvidia's dominant position in AI inference will depend on whether its performance claims can be verified in independent benchmarks—and whether CS-4 can be delivered quickly on a sufficient scale to meet the increasingly urgent high-speed token generation needs of hyperscale data center operators.
Cerebras' cloud inference business achieved 287% year-on-year growth in the second quarter, reaching $127.7 million in core cloud revenue. The launch of CS-4 marks a new phase in Cerebras' business model transformation from selling hardware to leasing reasoning capabilities.