Refacto AI

Industry story

OpenAI's Custom Inference Chip 'Jalapeño' Outperforms Nvidia Blackwell

cloud-costs gpu-supply inference model-pricing

OpenAI has publicly unveiled 'Jalapeño,' a custom AI inference chip developed in partnership with Broadcom, announced at the Hot Chips conference. Designed exclusively for large language model (LLM) inference — the process of running a trained AI model to generate outputs — the chip went from initial team hiring to manufacturing tape-out in approximately 16 months, an unusually fast development cycle for a custom ASIC (application-specific integrated circuit). Independent benchmarking firm SemiAnalysis tested the chip in OpenAI's lab using their InferenceX suite and found Jalapeño outperforms every Nvidia, AMD, and Google chip they've tested on multiple open-source models, measured by tokens generated per megawatt of power. The chip uses HBM4 memory (the latest high-bandwidth memory standard), achieves over 700 tokens per second per user on DeepSeek R1 at low concurrency, and does so without speculative decoding — a technique other chips rely on to boost speed — suggesting significant room for further improvement once that capability is added.

Analysis

Showing the shorter version.

OpenAI's Jalapeño Chip: What It Actually Changes

OpenAI showed off a custom inference chip called Jalapeño, built with Broadcom, and SemiAnalysis says it beats every NVIDIA, AMD, and Google chip they've tested on tokens per megawatt. The claim landing on every builder's desk: NVIDIA's inference moat is thinner than its training moat, and a lab can now prove it in one product generation.

Take the benchmark with caution. SemiAnalysis ran this in OpenAI's lab, on OpenAI's boxes, on workloads OpenAI picked. Tokens per megawatt is the metric you reach for when it flatters your architecture. Operators pay in tokens per dollar at production concurrency, and nobody showed that curve. A great benchmark in the builder's own basement is not the same as a chip that holds up when 10,000 users hit it at once.

The structural point survives even if the number is soft, though. A fabless lab closed the ASIC gap with NVIDIA in a single product generation, via Broadcom and TSMC, in 16 months. Custom ASIC cycles typically run three to four years from architecture to silicon. If that timeline holds under scrutiny, it's a methodology shift. Inference was always more attackable than training: it's a narrower, more predictable workload you can bake into silicon. Training still belongs to NVIDIA. But the inference moat just took a credible hit. AMD gets the worst of this, because its entire inference pitch was "cheaper NVIDIA," and a custom ASIC beats that story on both axes.

Nobody outside OpenAI can buy Jalapeño, so no external builder is changing hardware plans today. What changes is leverage. Every CTO renewing an H100 or H200 reservation next quarter now has a data point that the cost floor is not fixed. Push NVIDIA and your cloud on inference pricing now and see if the floor moves.

The flip side of cheaper OpenAI inference is concentration risk. A more vertically integrated OpenAI, running its own chip, its own model, and its own infrastructure, is less auditable and harder to exit. Cheaper and more locked-in are the same event.

The call: Before NVIDIA's next quarterly earnings call on 2026-11-18, NVIDIA will publicly emphasize a new or repriced inference-optimized product or offering in direct response to custom-ASIC pressure. Medium confidence. NVIDIA controls timing, but Jensen Huang has a long habit of answering competitive narratives fast rather than ceding the framing. An earnings call with this benchmark circulating is a setting where staying silent on inference price-performance would itself read as a concession.

The longer exposure for most builders is concentration. If inference silicon splinters across custom ASICs, the winners are the labs big enough to build their own. Everyone else is still renting, just from a shorter list of landlords.

Also covered this issue

Comments