ChipsAugust 25, 2026via TechCrunch AI
OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
Why it matters
OpenAI has shipped a custom inference chip (Jalapeño) that outperforms current SOTA on both tokens-per-user and power efficiency, signaling a shift toward vertical integration and compute economics at the inference layer — a critical lever for margin and latency in production AI workloads.
Key signals
- Custom chip named Jalapeño
- Benchmarked on Semianalysis InferenceX
- Higher tokens per user than current SOTA
- Superior throughput per kilowatt vs current SOTA
- Focus: fast inference at scale
- Implications for cloud compute economics and vertical integration
The hook
OpenAI's custom silicon just beat Nvidia on inference throughput per watt. Here's what it means for your cloud bill.
Tested on Semianalysis’s InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.