OpenAI: Jalapeño Chip Tops Efficiency Benchmarks
OpenAI Jalapeño delivers higher throughput per kilowatt and lower latency than rivals on GPT-OSS 120B, with deployment set for year-end.
SourceAnalysis
OpenAI released measured results for Jalapeño, its first custom inference chip, showing superior peak throughput per kilowatt and reduced token latency versus commercial systems on the InferenceX benchmark using GPT-OSS 120B, with strong results also on DeepSeek R1 and Kimi K2. The nine-month design-to-tapeout cycle relied on AI tools for iteration and arithmetic optimization. Deployment inside OpenAI infrastructure begins by year-end as part of a three-generation roadmap already advancing Gen 2 and Gen 3. OpenAI inference chip performance gains illustrate how efficiency expands total compute demand rather than reducing it.
Greg Brockman
@gdbPresident & Co-Founder of OpenAI