Analysis
OpenAI's Jalapeño Chip Wants to Rewrite the Economics of AI Inference

Image: Flickr / Wikimedia Commons / Unsplash

OpenAI's Jalapeño Chip Wants to Rewrite the Economics of AI Inference

First benchmark results claim big gains over Nvidia's Blackwell chips, but the comparisons are OpenAI's own, and the real payoff for API customers is still 12 to 18 months out.

August 25, 20268 min read

This article was produced by the AETW editorial team.

OpenAI published its first detailed benchmark results for Jalapeño, its custom AI inference chip built with Broadcom, claiming up to 4.1 times better performance per watt than Nvidia's Blackwell systems. The numbers are OpenAI's own, verification is partial, and the chip won't reach real production scale until 2027.

What OpenAI actually announced

At the Hot Chips conference on August 25, OpenAI presented its first detailed benchmark results for Jalapeño, the company's first custom AI inference chip, developed in close collaboration with Broadcom. The headline claim is straightforward: this AI inference chip is now a real contender against Nvidia's current data center hardware, not just a side project OpenAI is running in parallel with its Nvidia purchases.

Tested across three large language models, OpenAI's own GPT-OSS 120B, DeepSeek's R1 670B, and Moonshot AI's Kimi K2.5 1T, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the Nvidia GB200 and GB300 systems it was compared against. For highly interactive, multi-step workloads like AI agents, where delays compound across an entire task, OpenAI reported the gap widening to 2.1 to 4.1 times.

"The bottom line is that the results show a very, very significant performance advance over state of the art," Richard Ho, OpenAI's head of hardware, told reporters on a press call. "Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It's very efficient to serve a lot of customers, but it can also be very low latency."

These are OpenAI's numbers, with real caveats attached

OpenAI tested Jalapeño on InferenceX, an independent benchmark platform built by SemiAnalysis that measures the full process of serving an AI request. That framing matters for credibility, but it comes with limits worth stating plainly. SemiAnalysis verified the benchmark runs in person, in OpenAI's own lab, but has not yet run its full independent InferenceX suite on the chip, and has not tested it against AgentX, the firm's preferred benchmark for long-context, multi-turn workloads that better reflect real production traffic than short, single-turn tests.

There's also a question of which Nvidia hardware Jalapeño was measured against. The comparison used Nvidia's GB200 and GB300 Blackwell systems, which SemiAnalysis itself called somewhat incomplete, since Nvidia's newer Vera Rubin platform, which also uses HBM4 memory, has already begun shipping to customers and is the more relevant near-term rival. On throughput per megawatt, SemiAnalysis said Jalapeño's numbers still edge out Vera Rubin's published July results, though the two come out close to even on cost per token once total cost of ownership is factored in.

SemiAnalysis also pushed back on a narrower assumption that had circulated since Jalapeño was first unveiled in June: that this is a narrow accelerator built around OpenAI's own models the way a game console is built around its own games. The firm called it a generalized inference chip instead, noting it beat every Nvidia, AMD, and Google chip SemiAnalysis has been able to test on multiple open-source models, and did so without multi-token prediction, a technique the comparison chips were using to their advantage.

Why every AI lab wants its own inference chip now

Jalapeño's gains come from a full-stack design bet. OpenAI and Broadcom built the chip, memory, and networking together around real language-model workloads rather than adapting a general-purpose GPU after the fact. According to OpenAI, that let the team target the specific phases of inference that usually cause friction: prefill, which is compute-heavy, and the communication delays that happen when data has to move between chips. By keeping model state, including the KV cache used while generating a response, local to where it's needed, OpenAI says it minimized the data movement that otherwise idles compute while it waits.

That approach places Jalapeño inside a broader pattern in the AI chip market. Google has its TPU line, Amazon has been pushing its own Trainium and Inferentia silicon, and Anthropic has confirmed plans to build its own chips as well. Frontier AI labs increasingly want to own more of the hardware layer instead of renting it entirely from Nvidia, both to fit chips to their own workloads and to reduce exposure to Nvidia's pricing power and allocation decisions.

None of this replaces Nvidia in the near term. Jalapeño is an inference-only chip and cannot train models, so OpenAI's training compute still runs on Nvidia GPUs, and the company says it will keep deploying Nvidia hardware for both training and inference alongside its own silicon. Deployment is also just starting: OpenAI plans to bring Jalapeño into its own infrastructure in small volumes by the end of 2026, with meaningful scale arriving through 2027, and a second and third generation already in development.

Sources for this section

What changes for US builders paying for inference

For a US founder or small team building on OpenAI's API today, none of this changes anything yet. These are lab benchmarks on a chip that ships in small volumes late this year and doesn't reach meaningful production scale until sometime in 2027. Even once Jalapeño is running at scale inside OpenAI's own infrastructure, there's no guarantee the resulting efficiency gains show up as lower API prices rather than wider margins or reinvestment into more training compute. The realistic timeline for anything an outside developer would notice is closer to 12 to 18 months out, not the next product update.

Latency is the more interesting lever to watch specifically for agentic products. Multi-step agents compound delays across a task, so a chip that cuts end-to-end latency without sacrificing throughput is more likely to surface first in OpenAI's agent-heavy surfaces, like Codex and API-driven agent tooling, before it shows up as an across-the-board price cut on standard chat completions. The practical move for US operators is to track OpenAI's API pricing page and rate limits over the coming quarters rather than reading too much into the chip roadmap itself.

There's a second-order effect worth watching too. Every developer who has felt a GPU shortage translate into waitlists, capacity limits, or price increases elsewhere in the industry has a stake in whether OpenAI's custom silicon actually reduces its dependence on Nvidia allocation. If Jalapeño holds up anywhere close to its published numbers once it's running real production traffic, it's one more data point that inference compute is becoming less of a single-vendor bottleneck, which matters for pricing stability across the whole AI accelerator chip market, not just for OpenAI's own margins.

Sources for this section

What's still unproven

  • Whether SemiAnalysis's full independent InferenceX and AgentX runs confirm OpenAI's numbers once tested at arm's length.
  • Whether Jalapeño's efficiency edge holds up against Nvidia's newer Vera Rubin platform, rather than the older Blackwell generation it was benchmarked against.
  • Whether OpenAI passes any inference savings through to API customers as lower prices, or keeps the margin for itself.
  • How quickly the 2026 to 2027 deployment ramp actually reaches production scale, and whether real-world kernels hit the same upper-bound performance OpenAI is reporting from the lab.

Sources

Brian Weerasinghe

Founder and Editor

Brian Weerasinghe is the Founder and Editor of AI Eating The World. AI Eating The World is the independent AI publication for builders, operators, and leaders navigating how AI is changing work and the world.

Brian Weerasinghe is the founder and editor of AI Eating The World, where he covers artificial intelligence, tech companies, layoffs, startups, and the future of work. His reporting focuses on how AI is transforming businesses, products, and the global workforce. He writes about major developments across the AI industry, from enterprise adoption and funding trends to the real-world impact of automation and emerging technologies.

Trusted AI LeaderTrusted AI LeaderTrusted AI LeaderTrusted AI Leader
Trusted by 10,000+ builders

The AI brief for builders, operators, and leaders

Follow the AI developments reshaping work and the world, with practical context for what to do next.

Free, no spam, unsubscribe anytime.