Jalapeño’s first results show industry-leading speed and efficiency in AI inference
OpenAI reports its custom inference chip Jalapeño delivers up to 1.9x higher throughput per watt and 3.6x lower latency than leading systems across major AI models, enabling faster responses and more efficient agentic workloads.
OpenAI has completed initial testing of Jalapeño, its first custom inference chip, demonstrating a measurable advance in performance. The chip achieves higher throughput and lower latency simultaneously, a combination that typically requires tradeoffs in existing hardware. For customers, this means faster AI responses, more responsive agents, and improved reliability as demand increases. The gains are designed to make advanced AI more affordable and widely accessible while supporting OpenAI’s mission of benefiting humanity through artificial general intelligence.
Jalapeño’s performance was evaluated using InferenceX, a public benchmark from SemiAnalysis, across three major AI models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. Across these models, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than comparison systems. For highly interactive workloads, performance gains reached 2.1 to 4.1 times higher throughput. The chip’s measured sustained power remained at or below 550 watts despite a rated 700 watts, indicating strong efficiency.
The chip’s architecture was co-designed with OpenAI’s models to optimize for real-world language model workloads, addressing distinct phases like prefill and decode. Jalapeño minimizes data movement and communication delays by keeping model state local and integrating a high-bandwidth network to maintain speed and efficiency. This balanced design supports changing model architectures and excels in both compute-intensive and memory-constrained phases, particularly for agentic workloads where delays can compound across sequential tasks.
AI played a direct role in Jalapeño’s development, accelerating design cycles and optimizing arithmetic circuits to meet performance targets. The chip’s predictable programming model enables both human engineers and AI to efficiently map and schedule workloads. OpenAI plans to deploy Jalapeño within its infrastructure by the end of the year, with Gen 2 and Gen 3 already in development. The results underscore the benefits of a full-stack approach, where hardware, software, and models are designed together to deliver more responsive, capable, and efficient AI systems.