Refacto AI

Industry story

Amazon triples Nvidia GPU order to 3M chips for AWS

big-tech build-vs-buy cloud-costs gpu-supply reliability

Amazon just tripled its Nvidia GPU order to 3 million chips across 2027 and 2028, covering Blackwell Ultra, Rubin, and Rubin Ultra, plus Nvidia networking, CPUs, models, and robotics baked into AWS. The press release reads like conviction; the supply chain does not. CoWoS advanced packaging at TSMC is the real constraint, and no letter of intent builds more of it. The deeper problem is strategic: Amazon is handing Jensen Huang its largest order yet while simultaneously spending billions to replace him with Trainium, and those two positions cannot both be true.

Analysis

Showing the shorter version.

Three million GPUs "over two years" is a press release number. The real constraint has never been Amazon's appetite. It's CoWoS advanced packaging at TSMC, and no purchase order expands that capacity. Microsoft, Google, Oracle, and OpenAI-adjacent orders are all ahead of or alongside Amazon's in that queue. Hyperscaler GPU commitments have a long track record of being announced at round headline numbers and landing lower.

The part that doesn't add up: Amazon is tripling its order from the one supplier it's spending billions to replace with Trainium. You can't wave internal substitution at a vendor while handing them your biggest contract yet. Jensen Huang reads that the same way anyone does.

The roadmap commitment is the more interesting signal. Blackwell Ultra to Rubin to Rubin Ultra is a three-generation forward lock, and nobody does that on a spot demand blip. AWS internal forecasts are running ahead of public guidance. The co-optimization angle matters too: integrating Nvidia's full stack inside AWS generates joint tuning data that makes the pairing stickier every quarter. The 3M headline is the down payment; the 2028 architecture commitment is the durable part.

For teams building on AWS now, the practical tension is Trainium versus the Nvidia fabric. Above roughly 10,000 GPUs, the binding constraint shifts from raw compute to interconnect bandwidth. Nvidia's NVLink and Infiniband becoming the default fabric inside AWS clusters means distributed training jobs are gated by topology. If that fabric wins the high-throughput path, teams that extended their Trainium pilots for cost reasons will find themselves off the road of least resistance eighteen months from now. The failure mode is silent: you keep the pilot running because it feels safe, and the architecture decision gets made by the AWS console defaults.

There's a concentration risk nobody has priced yet. AWS and Nvidia are now co-integrated across networking, model families (Nemotron), and robotics (Omniverse, Isaac, Cosmos, Jetson). A CoWoS disruption or a chip-export escalation stops being a vendor incident and becomes an infrastructure event with blast radius across every startup renting AWS inference.

The call: Amazon will not disclose delivery of anywhere near 3 million Nvidia GPUs on the announced timeline. By Nvidia's Q4 FY2028 earnings (early 2028), AWS-bound Blackwell Ultra and Rubin shipments will have landed materially short of the announced pace, with TSMC packaging capacity named as the gating factor. The order is real. The quantity and timeline are aspirational.

One number to watch before then: reserved inference pricing on AWS in 2026. If Amazon's forward premium shows up as higher reserved rates, that's the skeptic's read confirmed early, and it lands on your inference bill before a single Rubin chip ships.

Also covered this issue

Comments