HeadlinesBriefing favicon HeadlinesBriefing.com

OpenAI Jalapeño Chip Delivers 1.5-4.1x AI Inference Gains

OpenAI Blog •
×

OpenAI has released first performance results for Jalapeño, its first custom inference chip, demonstrating industry-leading speed and efficiency. The chip delivers 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency compared to leading commercially available systems. For highly interactive workloads, Jalapeño achieved 2.1 to 4.1 times higher performance.

Testing across GPT‑OSS 120B, Deep Seek R1, and Kimi K2.5 1T models showed consistent gains for architectures developed both inside and outside OpenAI. On the largest model, Kimi K2.5 1T, Jalapeño delivered approximately 1.5 times higher peak performance per watt and 3.4 times lower latency. The chip is rated at 700 watts but sustained at or below 550 watts during testing.

OpenAI emphasizes a full-stack advantage, designing models, serving software, chips, memory, networking, and systems together using real workload data. Jalapeño represents the beginning of a multigenerational platform, with plans to ramp production to deliver faster, more capable, and efficient products. The company measured performance using the public Inference X benchmark from Semi Analysis, evaluating at matched user experience across operating ranges.