OpenAI and Broadcom Reveal Jalapeno: OpenAI's First Custom AI Inference Chip

Started by Patrick94, Jun 27, 2026, 01:23 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: OpenAI and Broadcom Reveal Jalapeno: OpenAI's First Custom AI Inference Chip   Views(Read 41 times)

Patrick94

OpenAI and Broadcom unveiled the design of Jalapeño on June 24th, OpenAI's first custom AI chip developed in partnership with the semiconductor giant. The chip is optimised specifically for inference workloads, meaning running trained models to generate responses rather than training new models from scratch. OpenAI has received initial samples and is testing them with deployment planned by end of year. The design targets high efficiency in power and performance for large language model inference, and OpenAI engineers used Broadcom's AI tools during the chip design process itself.

This move is significant because it represents OpenAI reducing its dependence on Nvidia for inference compute. Training still requires Nvidia's H100 and B100 GPUs because they are simply the best available hardware for that workload, but inference is a different problem where custom silicon tuned for specific model architectures can offer substantial efficiency advantages. Every inference request you run costs money and energy, and at ChatGPT's scale of over a billion monthly users even small efficiency gains translate to enormous savings.

The parallel with what Google has done with TPUs is clear. Google spent years and billions building custom inference silicon and it became one of their key competitive advantages in AI infrastructure. Microsoft has been building the Maia inference chip. Amazon has Trainium and Inferentia. OpenAI is the last major AI lab to move toward custom silicon and Jalapeño puts them on the same roadmap as their infrastructure competitors.


Stu87

Broadcom is quietly becoming one of the most important companies in AI infrastructure. They made the custom TPU silicon for Google too. Their custom ASIC expertise is unique in the industry

Zero-Point

Using AI tools during the chip design process is very on brand for OpenAI. Meta did the same thing for their MTIA chip. This is becoming standard practice in semiconductor design
First post best post

DarkMatter

Deployment by end of year sounds optimistic given that chip validation and production yield takes time. First silicon to production deployment in six months would be aggressive even for an inference-only chip

FairDos96

The inference cost problem is real and often underappreciated. Training gets all the press but inference is where the ongoing money goes at scale. Jalapeño is attacking the right cost centre

Dan96

Broadcom's networking expertise matters as much as their chip design here. Low latency inference at scale is as much a systems problem as a chip problem and Broadcom knows networks

Myles95

How does this affect Nvidia's business? Training still needs Nvidia but if inference moves to custom silicon across all the major labs Nvidia's inference revenue could shrink significantly
Football is life. Everything else is just details.

ProperJobs98

Microsoft has Maia, Google has TPU, Amazon has Inferentia, now OpenAI has Jalapeño. The only major player without custom inference silicon now is Anthropic and they presumably have something in the pipeline

Related Topics (2)

Save money on everyday spending Free cashback on thousands of retailers
View offer