How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip
OpenAI's Jalapeño AI accelerator chip, designed with LLM assistance in under 20 months, delivers 13.4 petaflops of 4-bit compute and claims up to 3.6x latency improvement over Nvidia's GB300.
On 25 August, OpenAI fully unveiled Jalapeño, the company’s debut AI accelerator chip. Jalapeño delivers up to 13.4 petaflops of 4-bit compute and accesses 232 gigabytes of the most advanced memory available, linking to it at a blazing 15.4 terabytes per second. Benchmarks cited by OpenAI show that Jalapeño can reduce end-to-end latency (the time between prompt to last token) by up to 3.6x when compared to Nvidia’s GB300—a chip the company currently relies on—and do so while consuming less power. Whether these figures translate into real-world gains once Jalapeño enters widespread service in OpenAI’s inference fleet remains to be seen, but performance is only half the story.
The other half is how the chip was designed—a process which, as you might expect, was accelerated by OpenAI’s large language models (LLMs). Jalapeño moved from first architecture concept to first silicon in under 20 months. Only nine months separated the first RTL—the register-transfer level code defining the chip’s logic—from tapeout, when the finished design goes to manufacturing. That’s a rapid timeline, yet experts believe it could soon look slow as LLMs improve and become more deeply integrated into chip…
- spectrum.ieee.orgHow OpenAI Used Its Own LLMs to Design Its Jalapeño Chipprimary