Industry story
OpenAI's Custom Inference Chip 'Jalapeño' Outperforms Nvidia Blackwell
cloud-costs gpu-supply inference model-pricing
OpenAI has publicly unveiled 'Jalapeño,' a custom AI inference chip developed in partnership with Broadcom, announced at the Hot Chips conference. Designed exclusively for large language model (LLM) inference — the process of running a trained AI model to generate outputs — the chip went from initial team hiring to manufacturing tape-out in approximately 16 months, an unusually fast development cycle for a custom ASIC (application-specific integrated circuit). Independent benchmarking firm SemiAnalysis tested the chip in OpenAI's lab using their InferenceX suite and found Jalapeño outperforms every Nvidia, AMD, and Google chip they've tested on multiple open-source models, measured by tokens generated per megawatt of power. The chip uses HBM4 memory (the latest high-bandwidth memory standard), achieves over 700 tokens per second per user on DeepSeek R1 at low concurrency, and does so without speculative decoding — a technique other chips rely on to boost speed — suggesting significant room for further improvement once that capability is added.
Analysis
Showing the shorter version.
OpenAI's Jalapeño Chip: What It Actually Changes
OpenAI showed off a custom inference chip called Jalapeño, built with Broadcom, and SemiAnalysis says it beats every NVIDIA, AMD, and Google chip they've tested on tokens per megawatt. The claim landing on every builder's desk: NVIDIA's inference moat is thinner than its training moat, and a lab can now prove it in one product generation.
Take the benchmark with caution. SemiAnalysis ran this in OpenAI's lab, on OpenAI's boxes, on workloads OpenAI picked. Tokens per megawatt is the metric you reach for when it flatters your architecture. Operators pay in tokens per dollar at production concurrency, and nobody showed that curve. A great benchmark in the builder's own basement is not the same as a chip that holds up when 10,000 users hit it at once.
The structural point survives even if the number is soft, though. A fabless lab closed the ASIC gap with NVIDIA in a single product generation, via Broadcom and TSMC, in 16 months. Custom ASIC cycles typically run three to four years from architecture to silicon. If that timeline holds under scrutiny, it's a methodology shift. Inference was always more attackable than training: it's a narrower, more predictable workload you can bake into silicon. Training still belongs to NVIDIA. But the inference moat just took a credible hit. AMD gets the worst of this, because its entire inference pitch was "cheaper NVIDIA," and a custom ASIC beats that story on both axes.
Nobody outside OpenAI can buy Jalapeño, so no external builder is changing hardware plans today. What changes is leverage. Every CTO renewing an H100 or H200 reservation next quarter now has a data point that the cost floor is not fixed. Push NVIDIA and your cloud on inference pricing now and see if the floor moves.
The flip side of cheaper OpenAI inference is concentration risk. A more vertically integrated OpenAI, running its own chip, its own model, and its own infrastructure, is less auditable and harder to exit. Cheaper and more locked-in are the same event.
The call: Before NVIDIA's next quarterly earnings call on 2026-11-18, NVIDIA will publicly emphasize a new or repriced inference-optimized product or offering in direct response to custom-ASIC pressure. Medium confidence. NVIDIA controls timing, but Jensen Huang has a long habit of answering competitive narratives fast rather than ceding the framing. An earnings call with this benchmark circulating is a setting where staying silent on inference price-performance would itself read as a concession.
The longer exposure for most builders is concentration. If inference silicon splinters across custom ASICs, the winners are the labs big enough to build their own. Everyone else is still renting, just from a shorter list of landlords.
OpenAI showed off a custom inference chip called Jalapeño, built with Broadcom, and SemiAnalysis says it beats every NVIDIA, AMD, and Google chip they've tested on tokens per megawatt. The claim landing on every builder's desk this week: NVIDIA's inference moat is thinner than the training moat, and a lab can now prove it in one product generation. What does that mean for anyone running models in production, and who should actually change plans?
This is a briefing, so the frame is simple. Reversibility is Type 2 for most readers. Nobody outside OpenAI can buy Jalapeño, so no one is committing to anything today. The real decision hiding underneath is whether you keep treating NVIDIA pricing as fixed when you plan next year's inference budget. The forcing function is not a chip you can rack. It's the negotiating leverage this announcement hands to every large inference buyer, right now.
The Skeptic. SemiAnalysis ran this in OpenAI's lab, on OpenAI's boxes, on workloads OpenAI picked. That is a press tour with a spreadsheet, not independent benchmarking. Tokens per megawatt is the metric you reach for when it flatters your architecture. Operators pay in tokens per dollar at production concurrency, and nobody showed that curve. The "no speculative decoding yet" line is unfalsifiable headroom dressed as modesty. Google's TPUs also look mediocre on slides and quietly run most of the world's inference. For a PM: a great benchmark in the builder's own basement is not the same as a chip that holds up when 10,000 users hit it at once.
The Compute Pragmatist. Strip the hype and one real thing remains: a fabless lab closed the ASIC gap with NVIDIA in a single generation, via Broadcom and TSMC, on 16 months. NVIDIA's inference moat was always softer than its training moat, because inference is a narrower, more predictable workload you can bake into silicon. Training still belongs to NVIDIA entirely. But HBM4 is the choke point now. Whoever gets SK Hynix and TSMC allocation controls the next inference generation, and OpenAI just jumped the queue. AMD takes the worst of this. Its whole inference pitch was "cheaper NVIDIA," and a custom ASIC beats that story on both axes. In plain terms: the thing that runs your model may stop being a GPU you rent from one vendor.
The Researcher. The 16-month tape-out is what people are sleeping on. Custom ASIC cycles run three to four years from architecture to silicon. If 16 months holds under scrutiny, that is a methodology shift, not just a fast team. HBM4 this early matters too, since most production silicon is still on HBM3e. Framing the win without Multi Token Prediction, where every competing chip used its best MTP config, makes the efficiency gap look conservative rather than cherry-picked. The open question that decides everything: does this architecture generalize past autoregressive decode, or is it a narrow bet on one workload class? For a PM: they optimized hard for one specific job, and we don't yet know if it does anything else well.
The Enterprise Buyer. A chip I can't buy doesn't change my procurement, but it changes my leverage. Every CTO renewing an H100 or H200 reservation next quarter now has a data point that the cost floor is not fixed. That's real. The flip side is concentration risk. If OpenAI runs its own chip, its own model, and its own infrastructure, my vendor is now more vertically integrated and less substitutable. Fewer places to audit, fewer places to exit. For a PM: cheaper inference from OpenAI is good until you realize you're renting from a landlord who now owns the building, the plumbing, and the electric company.
Three tensions worth naming. The Skeptic says the benchmark is marketing; the Compute Pragmatist says the structural point survives even if the number is soft, because a fabless lab shipping a competitive inference ASIC at all is the news. They can both be right. The chip may not beat Blackwell in your rack, and NVIDIA's inference pricing power still just took a credible hit.
Second, the Researcher's optimism about the 16-month cycle collides with the Skeptic's point that tape-out is not production. A tape-out is a chip that came back from the fab, not a fleet running at scale. The gap between those two is where most silicon dreams die.
Third, the Enterprise Buyer wants the price relief and fears the lock-in that comes with it. Cheaper OpenAI inference and a more vertically integrated, less auditable OpenAI are the same event.
So what does this actually hinge on? One belief: does credible custom-silicon competition force NVIDIA to move on inference pricing before it forces anything on availability? The chip stays internal, so external builders get no hardware this cycle. What they get, if anything, is leverage. If you run large inference, the thing to test is not Jalapeño. It's your next reserved-instance quote. Push NVIDIA and your cloud on inference pricing now and see if the floor moves. Design the load test you'd run on custom silicon anyway, so you're ready if OpenAI ever opens it or a rival ASIC ships.
The council leans toward the structural read over the benchmark. The tokens-per-megawatt number may not survive contact with production. The signal that inference silicon is now contestable already survived, the moment SemiAnalysis put OpenAI on the same chart as Blackwell and OpenAI won on the axis it chose.
Prediction: Before NVIDIA's next quarterly earnings call on 2026-11-18, NVIDIA will publicly emphasize a new or repriced inference-optimized product or offering (a Blackwell inference SKU, a rack config, or explicit inference price-performance claims) in direct response to custom-ASIC pressure, rather than let the Jalapeño narrative stand unanswered.
Confidence: Medium. Clear incentive and a fixed earnings anchor, but NVIDIA controls timing.
Why: SemiAnalysis put OpenAI's chip above Blackwell on tokens per megawatt, and that framing attacks NVIDIA exactly where it's weakest, since inference is a narrower workload that custom silicon can target while training stays NVIDIA's. NVIDIA's entire data-center revenue mix depends on customers believing GPUs are the default for both training and inference, and Jensen Huang has a long habit of answering competitive narratives fast and loudly rather than ceding the framing. An earnings call with this benchmark circulating is a setting where staying silent on inference price-performance would itself read as a concession, which is why the counter is more likely than quiet. The opposite outcome, NVIDIA ignoring it entirely, would break with how Huang has handled every prior custom-silicon and TPU threat.
Revisit by 2026-11-18: We're right if NVIDIA promotes an inference-specific product, config, or price-performance claim on or before its November earnings call in a way that reads as answering custom-ASIC pressure. We're wrong if NVIDIA makes no inference-specific competitive move and leaves the Jalapeño comparison unaddressed.
The one thing worth adding: the real exposure for most builders is concentration. If inference silicon splinters across custom ASICs, the winners are the labs big enough to build their own. Everyone renting is still renting, just from a shorter list of landlords.
Also covered this issue
-
OpenAI announces 80% price cut with 'Luna' model, pledges ongoing efficiency gains
techcrunch-ai
OpenAI's 80% price cut on Luna forces engineering teams to choose between month-to-month flexibility and annual contracts that bet on unproven efficiency promises.
-
Mistral Partners with HUMAIN for Sovereign AI in Saudi Arabia
mistral-blog
Mistral's bet on sovereign-compute decoupling could let your team run frontier models on customer-owned infrastructure instead of hyperscaler lock-in, if the governance layer actually works.
Comments