- High-performance memory accounts for more than 75% of AI server hardware costs, according to Google Cloud’s Senior Director of Supply Chain Infrastructure, a constraint the TPU 8 family is explicitly designed around.
- Google’s TPU 8i inference chip carries 288GB of HBM and 384 MiB of on-chip SRAM, keeping key-value caches on-chip; the TPU 8t scales to 9,600-chip clusters with a 2-petabyte shared HBM pool for training.
- The TurboQuant algorithm compresses KV-caches 6x without model degradation, and a new Marvell agreement extends Google’s custom silicon to storage controllers, network interfaces and memory-interface components.
Memory, not compute, is now the dominant cost in AI server hardware, and at the SEMICON Taiwan Memory Summit on September 1, 2026, Google Cloud Senior Director of Supply Chain Infrastructure Nikhil Cherian put a number on it: high-performance memory exceeds 75% of AI server hardware costs. The 8th-generation TPU family, split into two purpose-built variants and paired with a lossless compression algorithm, is Google’s attempt to make that arithmetic less punishing.
Built Around the Memory Wall
The TPU 8i is Google’s inference chip, and its architecture reflects exactly where inference bottlenecks live. It carries 288GB of High Bandwidth Memory and 384 MiB of on-chip SRAM, enough to hold dynamic dialogue states and key-value caches entirely on-chip, cutting off-chip memory latency to zero for those operations. Key-value cache access is one of the primary latency drivers in large language model serving; keeping it on-chip rather than shuttling data off to external HBM changes the response-time profile considerably.
The TPU 8t handles training at a different scale. Clusters of up to 9,600 chips form a shared HBM pool reaching 2 petabytes, with TPU Direct Storage technology designed to manage data movement across that pool. Both chips sit within Google’s unified stack, combining custom hardware with open software and flexible consumption models.
TurboQuant and the Software Layer
Hardware headroom alone does not close the memory gap. Google’s TurboQuant algorithm attacks the same problem from the software side, using lossless quantization to compress key-value caches 6x and reduce memory usage without degrading model output. As mixture-of-experts and multimodal architectures scale up, AI workloads shift from compute-constrained to memory-constrained. A 6x reduction in KV-cache footprint is worth considerably more in that environment than it would have been two years ago.
Beyond TPUs: The Marvell Deal
Google’s custom silicon ambitions run deeper than the TPU chips themselves. An August 2026 agreement with Marvell Technology initially disclosed through a securities filing, covers AI inference accelerators alongside storage controllers, network interface controllers, memory-interface controllers and near-memory computing technologies. The warrant structure ties equity awards to up to $120 billion in cumulative revenue through Marvell’s fiscal 2033, a figure that points to long-term supply alignment rather than a transactional partnership.
The logic is direct: if memory dominates hardware costs, optimising compute alone leaves most of the cost curve untouched. By extending custom silicon into storage, networking and memory interfaces, Google is working the problem at the system level. Broadcom remains a separate key partner for future TPU generations, so the Marvell deal broadens Google’s supplier base rather than replacing an existing relationship.
TPU Revenue Projections
Morgan Stanley has flagged the TPU opportunity as significant, with projections pointing to a substantial increase in TPU-related Google Cloud revenue, though those figures assume continued enterprise uptake at a pace that remains to be demonstrated. The memory-cost problem Cherian described at SEMICON Taiwan is not unique to Google, every hyperscaler building large-scale inference infrastructure faces the same arithmetic. What Google is betting on is that a tightly integrated hardware-software stack, from TPU silicon to KV-cache compression to custom storage controllers, compounds efficiency gains in a way that general-purpose infrastructure cannot match. Whether the TPU 8i and 8t bear that out in production workloads is the real test.



