Industry story
Google TPU 8i delivers 80% better inference perf/dollar for agents
ai-in-adtech cloud-costs engineering
Full analysis
Google announced its eighth-generation TPUs at Cloud Next '26, including the TPU 8i optimized for inference. TPU 8i uses a new 'Boardfly' topology to connect 1,152 TPUs in a single pod, features 3x more on-chip SRAM than prior generations to keep larger KV caches (the memory buffer used during model inference) entirely on-chip, and includes a specialized Collectives Acceleration Engine. Google claims 80% better performance per dollar for inference versus the prior generation — directly citing the goal of enabling 'millions of concurrent agents to run cost-effectively.'
Comments