The Zhitong Finance App learned that Morgan Stanley released a semiconductor industry research report stating that the shortage of DRAM memory will continue through the current AI industry cycle, and the expansion of AI computing power will not wait for new fabs to be completed and put into operation. The industry is circumventing memory bottlenecks through the three major technology paths of hardware downsizing, decoupling inference architectures, and CXL memory pooling, and continues to advance computing power construction under supply constraints. The bank continues to be optimistic about storage leaders such as Micron (MU.US) and SNDK.US (SNDK.US), and is optimistic about incremental opportunities in CXL connectivity and heterogeneous reasoning tracks. It recommends Astera Labs (ALAB.US), Marvell (MRVL.US), Cerebras (CBRS.US), and NVDA.US (NVDA.US).
Damo emphasized that the current shortage of AI memory is not a short-term disturbance, but rather a structural contradiction where the increase in computing power performance far exceeds the speed of memory supply. The scale of cutting-edge large-scale models doubles every 6 months, the context window of the head model grows 5-6 times a year, the demand for superimposed inference and concurrency continues to increase, and demand for memory capacity and bandwidth continues to explode. However, the DRAM fab construction cycle is several years long, supply elasticity is extremely low, and the shortage pattern will continue for several years. Nvidia CEO Hwang In-hoon also publicly stated that the industry needs to change its thinking and deal with memory constraints through architectural innovation rather than simply waiting for production capacity to expand.
Facing the tight supply of HBM and main memory, the industry's most direct response is to selectively reduce single-device memory specifications (de-speccing). Taking Nvidia's Rubin architecture as an example, the LPDDR5 capacity of a single rack was reduced from 54 TB to 28 TB, and the HBM capacity of a single GPU was reduced from 288 GB to 192 GB to guarantee the overall shipment volume by reducing the height of the memory stack. Damo pointed out that downsizing does not eliminate memory bottlenecks, but rather shifts pressure between memory levels: after cutting high-speed local memory, non-high-frequency data such as KV caches sink to NAND storage, while cross-GPU data interaction increases, driving demand for network interconnection bandwidth. Essentially, it is a faster network connection and lower cost storage resources to supplement scarce high-bandwidth memory.
The second path to breaking the game is inference architecture disaggregation (disaggregation). AI inference includes two stages with significant differences between prefill (prefill) and decode (decode): Prefill mainly consumes computing power, while Decode is highly dependent on memory bandwidth. Traditional architectures use the same accelerator to balance both types of tasks, and resource utilization is low. Currently, the industry is accelerating towards heterogeneous reasoning, splitting the two stages into different hardware: computational power-intensive prefills are handled by general-purpose GPUs, and memory-bandwidth-intensive Decodes are processed by dedicated accelerators equipped with high-capacity on-chip SRAM.
Typical examples include Cerebras' wafer-level engine and Nvidia's acquisition of the Groq LPU architecture, which can provide memory bandwidth efficiency far beyond traditional GPUs during the decoding phase. Damo believes that heterogeneous reasoning will become an important evolutionary direction for AI infrastructure, and computing power vendors that specialize in decoding will gain a clear incremental market.
The third path is to restructure the memory architecture with CXL technology. CXL (High Speed Computing Interconnect) enables memory expansion, sharing, and pooling through high-speed interconnection, so that memory is no longer bound to a single processor, and has become a core technology path to break through memory capacity constraints. According to Damo estimates, AI demand will drive the CXL and related memory add-on chip market to reach about 6 billion US dollars by 2030, far exceeding the market volume of traditional CPU memory expansion. Its core value is to build a hierarchical memory architecture: the highest frequency data is kept in HBM, sub-high frequency data is placed in the CXL memory pool, low frequency cold data sinks to NAND, and the cost gradient is used to match the data access frequency.
On the main line of investment, Damo reiterated three major directions: one is to continue to surpass Micron and SanDisk, and the downsizing is due to a shortage of supply rather than weakening demand. The second is to lay out CXL and scale-up connectivity leaders Astera Labs and Marvell; third, to focus on Cerebras, the beneficiaries of heterogeneous reasoning, and Nvidia, which perfects the heterogeneous layout through the acquisition of Groq.
Risk warning: AI computing power demand is growing less than expected; implementation of CXL technology is slower than expected; memory capacity expansion exceeds expectations.