-+ 0.00%
-+ 0.00%
-+ 0.00%

In the age of AI reasoning, demand for TPU is exploding! Damo: Google's million TPUs leverage $200 billion in cloud computing revenue

Zhitongcaijing·08/25/2026 14:09:18
Listen to the news

The Zhitong Finance App learned that Wall Street financial giant Morgan Stanley's latest estimates show that as the AI computing power industry chain further shifts from centralized AI training to large-scale AI inference processes, the AI ASIC computing power cluster led by Google's TPU is being upgraded from cost reduction and efficiency improvement and customized planning tools within Google's cloud computing business to a high-margin AI cloud infrastructure business that can be sold abroad. The promise of one million TPUs and a potential large-scale batch supply of about 35 billion US dollars prompted the Morgan Stanley analyst team to drastically raise the Google TPU system sales price assumption from $20 billion to $27 billion per gigawatt, and raise the gross margin assumption from 20% to 30%.

According to a recent research report released by Morgan Stanley, the forecast is that Google's TPU related cloud revenue in 2027 and 2028 will reach 84 billion US dollars and 108 billion US dollars, respectively, for a total of close to 200 billion US dollars. This also means that Google's parent company Alphabet (GOOGL.US) may further transform from being the main bearer of huge AI capital expenses to a superplatform for leasing/selling AI computing power systems, and TPU may become another large-scale revenue growth engine after search advertising and public cloud computing IaaS+ SaaS businesses.

One million TPU leveraged $35 billion in orders! Damo: Google points to the $200 billion computing power gold mine

For mature models, open source models, and AI agent workloads that have a relatively stable model structure, huge call volume, and can maintain high utilization rates for a long time, special AI chips such as TPU (Google TPU is the most typical AI ASIC technology route developed by major cloud computing companies) can be deeply optimized around low-precision matrix computation, memory access, and chip interconnection, thereby improving the cost per token, throughput per watt, and total cost of ownership; at the same time, Google's TPU computing power clusters can also undertake large-scale training tasks, so the more accurate trend is GPU and ASIC /TPU forms a heterogeneous division of computing power.

In other words, AI training operator processes and cutting-edge AI workloads with the most complex architectures and the fastest changing rate still require AI GPU clusters — cutting-edge model pre-training, reinforcement learning, and rapidly changing new operators still rely more on GPU programmability, CUDA ecosystem, and NVLink/NVSwitch cluster capabilities, while large-scale AI inference workloads around mature and open source AI models, Copilots, and AI Agents (AI agents) are increasingly suitable for dedicated self-developed AI chips. Specifically, Google divided the eighth-generation TPU into a training-type TPU 8t and an inference TPU 8i:8i that optimizes KV Cache and low-latency inference through larger on-chip SRAM, 288GB HBM, and a dedicated pooled communication engine. Officially, its inference cost is up to 80% higher than Ironwood.

According to institutions such as Morgan Stanley and Wedbush Securities that continue to be optimistic about the investment prospects of the AI computing power industry chain, the almost endless cutting-edge computing power in the AI reasoning era and the computing power demand surrounding AI agents enabled AI ASIC to grow into an important component of the second trillion-level computing power ecosystem without destroying GPU demand — instead further strengthening the AI computing power industry chain investment logic that “the AI semiconductor supercycle is not a single GPU cycle, but an exponential increase in the silicon content of the entire data center.”

A team of analysts led by Morgan Stanley analyst Brian Novak wrote in a recent report to clients: “Notably, recent media reports show that Google, a subsidiary of Alphabet, promised to supply about 1 million TPUs at a price of about 35 billion US dollars — equivalent to about 1.3 gigawatts, according to our estimates — equivalent to about $27 billion per gigawatt after our subsequent increase. Previously, we assumed that in these first-party partnerships, Google would sell racks and systems at a price of $20 billion per gigawatt, with a gross margin of 20%.” ”

Furthermore, the further expanded customized chip development agreement between Google and Maywell Technology — our semiconductor research team wrote an analysis here — is consistent with the judgment that TPU pricing and average sales price are higher than our previous assumptions, and can be said to support this judgment. In light of this new information, we raised TPU's first-party sales forecast substantially to $27 billion per gigawatt... while gross sales margin increased substantially to 30%. Overall, we currently expect Google to sell 0.3 gigawatts of TPU systems in the second half of 2026, 3.2 gigawatts in 2027, and 4.2 gigawatts in 2028. According to this estimate, Google's cloud computing business revenue data related to TPU in 2027 and 2028 will reach 84 billion US dollars and 108 billion US dollars, respectively.” A team of analysts led by Brian Novak wrote.

The team of Daimo analysts led by Novak continued to give Google an “plus” rating, while the target price was reiterated at $400. As of the beginning of the US stock market on Tuesday, Google's parent company Alphabet's stock price hovered around $348 per share.

Google teamed up with Marvell, Microsoft Maia 200, and Anthropic to hunt TPU veterans. AI giants compete for “silicon rights” in the AI reasoning era

Token economics in the AI inference era is completely different from AI training workloads. Training includes forward propagation, backward propagation, gradient synchronization, and frequently changing research-based operators, and places more emphasis on versatility, software maturity, and cluster expansion capabilities; inference can be split into computationally intensive prefills (prefills) and memory bandwidth, KV caching, and latency-sensitive decoding. ASIC/XPU can remove unnecessary general-purpose circuits and harden FP8/FP4 low-precision matrix operations, attention mechanisms, mixed expert models (MoE), KV cache, and data movement paths, thereby improving the number of tokens per watt, number of tokens per dollar, and the certainty of service level agreements (SLAs).

It should be emphasized that although the model call mode of AI agents is more suitable for ASICs than GPUs, their planning, tool call, search enhanced generation (RAG), and state management still require collaboration between CPU, memory, network, and storage, so the intelligent era will further strengthen the “heterogeneous computing system.”

Microsoft (MSFT.US), one of the US tech giants, plans to launch its latest self-developed AI accelerator MAIA soon. In addition, the new agreement between Google and Fabless chip company Marvell (Maywell Technology) covers various customized chips such as AI accelerators, storage, and memory controllers, highlighting that the AI computing power infrastructure construction process led by AI application leaders and hyperscale cloud computing vendors is being completely upgraded from “collective procurement of Nvidia GPUs” to Nvidia GPU+AMD GPU+self-developed AI special integrated circuits (AI ASICs) Heterogeneous computing system with large-scale cooperative operation of /XPU+CPU/DPU. As the AI large model architecture gradually stabilized and the number of token calls in various industries around the world showed exponential growth, highly concurrent AI inference workloads became more and more suitable for reducing the cost per token through customized chips.

The new agreement between Google and Marvell (or Maywell Technology) fully covers customized chips such as AI accelerators, storage, and memory controllers. Potential purchases correspond to Marvell's maximum revenue of about 120 billion US dollars as of FY2033, essentially diversifying the dependence on Broadcom, Google's long-standing technical partner for TPU, a single Fabless design partner.

The Microsoft Maia 200 uses TSMC's 3-nanometer manufacturing process, FP8/FP4 tensor core, 272MB on-chip SRAM, and custom on-chip network (NoC) to reduce the inference costs of Azure, Copilot, and large models; its core advantage over Nvidia GPUs is not “any model is faster,” but rather that Microsoft can jointly optimize chips, compilers, models, data center networks, and cloud scheduling, and avoid paying all profit premiums to external GPU manufacturers.

According to a Microsoft research report, the Maia 200 has achieved a 30% increase in performance per dollar compared to the latest generation hardware in its fleet, and its self-developed MAI model has improved performance per watt by about 40% when running on the Maia 200. This is the most dangerous competitiveness of ASIC/XPU in the age of reasoning: when billions of similar token generation tasks can be highly standardized, the flexibility premium of Nvidia-led “universal GPUs” may not necessarily be worth paying for every inference request.