Industry story
Google TPU 8i delivers 80% better inference perf/dollar for agents
ai-in-adtech cloud-costs engineering
Google's TPU 8i isn't a chip announcement — it's an infrastructure bet that agentic AI workloads are about to scale hard and fast. The 80% inference performance-per-dollar gain over the prior generation, a 1,152-chip Boardfly pod, and 3x more on-chip SRAM are all pointed at one thing: running millions of concurrent agents without the economics falling apart. Every hyperscaler is making this bet right now, but Google is the only one building custom silicon end-to-end for it. Whether the agents actually show up at that scale is the question worth watching.
Full analysis
Google announced its eighth-generation TPUs at Cloud Next '26, including the TPU 8i optimized for inference. TPU 8i uses a new 'Boardfly' topology to connect 1,152 TPUs in a single pod, features 3x more on-chip SRAM than prior generations to keep larger KV caches (the memory buffer used during model inference) entirely on-chip, and includes a specialized Collectives Acceleration Engine. Google claims 80% better performance per dollar for inference versus the prior generation — directly citing the goal of enabling 'millions of concurrent agents to run cost-effectively.'
Comments