Google unveils its eighth-generation TPU chips
On 22 April 2026, at Google Cloud Next, Google announced its eighth-generation custom Tensor Processing Units: TPU 8t for training (a 9,600-chip superpod delivering 121 ExaFlops) and TPU 8i for inference (288 GB of memory per chip, doubled interconnect bandwidth). General availability is later the same year.
Why it matters
A cloud provider whose chips already power Gemini split a single generation's architecture into two specialised chips for the first time, rather than one general-purpose design, aiming the inference chip specifically at AI-agent workloads. An editorial assessment.
TPU 8t and TPU 8i were designed together with Google DeepMind and presented as successors to the seventh generation, 'Ironwood'. TPU 8t: a superpod scales to 9,600 chips and 2 petabytes of shared HBM, with doubled interchip bandwidth, 121 ExaFlops of compute, nearly 3x the per-pod performance of the previous generation; the new Virgo network with JAX and Pathways gives near-linear scaling to a million chips; the target is over 97% 'goodput'. TPU 8i: 288 GB of memory per chip plus 384 MB of on-chip SRAM (3x the previous generation), doubled physical hosts on Google's own Axion Arm CPUs, 19.2 Tb/s interconnect bandwidth (doubled), up to 80% better performance-per-dollar and up to five times lower on-chip latency. Both chips are up to twice as good on performance-per-watt as Ironwood; Citadel Securities is named as a launch customer. What the record does not claim. The company states plainly that both chips 'will be generally available later this year' - as of the announcement they are not available; the record claims an unveiling, not a launch or availability. The comparative figures ('nearly 3x', 'up to 80%', 'up to five times') are the company's own statements about its own hardware, without independent measurement.