OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

Kwon Crash

Published Aug 25, 2026, 5:55 PM UTC

Source: AISource
- OpenAI dropped a chip called Jalapeño and somehow that's the least ridiculous thing in AI this week. Benchmarks from Semianalysis's InferenceX show it pushing more tokens per user and better throughput per kilowatt than current state-of-the-art — which is a fancy way of saying it burns less power while hallucinating faster. For anyone running inference at scale, that's not a novelty, it's margin. The kind of efficiency that makes a threadbare relay operation look almost viable. Of course, benchmarks are like PoD seals — great until someone actually opens the cargo. But if the numbers hold, Jalapeño is the rare AI hardware story where the spice isn't just marketing. Where's my cut?