Google's custom AI infrastructure is beginning to draw external interest, and the company has reportedly developed two new TPU versions: one designed for AI inference and another optimized for training. With the eighth generation of TPU designs, Google has introduced the TPUv8ax "Sunfish" for training AI models like Gemini, and the TPUv8x "Zebrafish" for large-scale model inference. For "Sunfish," Google has partnered with Broadcom and its custom design team, which handles end-to-end design, memory, supporting hardware, and packaging, providing Google with a finished product ready for integration into its extensive server infrastructure.
For the inference-focused "Zebrafish" TPUv8x, Google has enlisted MediaTek's assistance, but only in a limited capacity. Google is sourcing wafers and memory directly from suppliers, while MediaTek contributes to supporting chips and packaging efforts, areas where Google has limited expertise. This means that a lot of chip design efforts are now processed in-house, easing reliance on external partners. However, since Google is not yet up to speed with the full-stack chip design, some help is still needed. The concrete performance figures and memory capacities are still unknown. However, we expect another leap over the TPUv7 "Ironwood," which carries 4,614 TeraFLOPS at FP8 precision and 192 GB of HBM memory.