Chips & Hardware
OpenAI unveils Jalapeño, its first custom inference chip
OpenAI has unveiled Jalapeño, its first custom-built inference processor developed with Broadcom, aiming to improve performance-per-watt and reduce reliance on Nvidia hardware.
On Wednesday, OpenAI unveiled Jalapeño, its first custom-built inference processor. In the context of artificial intelligence, an inference processor is a chip designed specifically for running pre-built AI models. Developed in collaboration with Broadcom, which served as a collaborator in the design and manufacturing of the custom chip, the new custom-built inference processor is still being tested. However, OpenAI reports that early results show significantly better performance-per-watt than current state-of-the-art alternatives.
The development of Jalapeño has long been rumored as a way for OpenAI to reduce its dependence on Nvidia, which is the current provider of GPUs that OpenAI aims to reduce dependence on. By designing its own silicon, OpenAI follows other technology companies like Google and Amazon, both of which have built custom chips for AI acceleration. These custom chips, often called AI accelerators, are silicon designed specifically to speed up machine learning workloads. While the new chip targets inference, performance-intensive tasks like pre-training—the initial, performance-intensive phase of training AI models—will likely still rely on Nvidia hardware.
The partnership with Broadcom was officially announced in October. Shortly after, OpenAI president Greg Brockman, who explained the chip development approach, noted that the company has a deep understanding of the workload. Brockman stated that they have been looking for specific workloads that are underserved, and asking how they can build something that will be able to accelerate what is possible. The company is already building agentic products like Codex, an agentic product built by OpenAI, and it views custom silicon as part of a broader strategy to control its entire infrastructure. According to OpenAI, the company is not only developing frontier models or building products on top of them, but is also designing the infrastructure underneath them. This infrastructure includes chip architecture, kernels, memory systems, networking, scheduling, deployment systems, and the product experience. “Because OpenAI operates across the stack, each layer can be optimized around the same goal: making its models faster, more reliable, and more affordable for users,” the company stated.
Why it matters
Developing custom silicon allows OpenAI to optimize its infrastructure from the chip level up, potentially lowering inference costs and reducing its dependence on Nvidia GPUs.