OpenAI's first custom inference chip Jalapeno beats existing hardware on speed and power efficiency

Started by Amy, Yesterday at 03:21 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: OpenAI's first custom inference chip Jalapeno beats existing hardware on speed and power efficiency   Views(Read 37 times)
Active members in this topic:
Amy(1) Reacher Erin(1)

Amy

OpenAI published its first performance results this week for Jalapeno, its first custom designed chip built specifically for AI inference rather than training, and the early numbers show a real, measurable advance in both speed and power efficiency compared to existing commercially available hardware. Tested across three different large language models, GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, Jalapeno delivered 1.5 to 1.9 times more AI work per watt at peak throughput and between 1.7 and 3.6 times lower end to end latency than the comparison systems it was benchmarked against.

The chip was tested using InferenceX, a public benchmark from SemiAnalysis that measures the complete process of actually serving a real AI request rather than just testing raw isolated compute throughput alone. OpenAI specifically framed its own preferred measurement standard as performance per unit of power rather than performance per individual chip, arguing that's the more especially useful comparison for customers actually paying for real compute at scale. Jalapeno is officially rated at 700 watts, though the company says its measured sustained power draw stayed at or below 550 watts across all the specific workloads actually tested during evaluation.

The underlying architecture was built specifically around how modern language model inference actually works in practice, recognizing that prefill, when a system first processes an incoming prompt, is primarily compute intensive, while decode, when the system then generates its response one token at a time, is instead constrained mainly by memory bandwidth rather than raw compute power. Jalapeno was designed to minimize data movement and communication delays between these two genuinely different phases, keeping model state like the KV cache used during response generation explicitly placed and kept local rather than needing to shuffle it constantly back and forth between separate physical chips.

OpenAI says AI itself played a direct, active role in Jalapeno's own development, helping the team move from initial design all the way through to tapeout in just nine months by exploring different implementations and continuously iterating faster on real workloads. In one specific concrete example, engineers used OpenAI's own Codex tool alongside an internal model called GPT-Astra to bring three additional open weight models that weren't originally part of Jalapeno's production plan up to truly high performance within just two months, with AI generated implementations for select attention and mixture of experts blocks running 1.5 to 1.8 times faster than the equivalent existing human expert written code. The company plans to begin deploying Jalapeno within its own compute infrastructure by the end of this year, with a second generation chip already deep in development and a third generation currently taking shape behind it

Normal is overrated

Reacher Erin

Using AI to help design the actual chip meant to run AI more efficiently is such a clean, almost recursive little loop once you actually stop and think it through properly. Really fascinating watching AI development start to directly bootstrap its own underlying hardware infrastructure like this in a fairly concrete, measurable way
Long time lurker, first time poster

Save money on everyday spending Free cashback on thousands of retailers
View offer