The full stack behind abundant intelligence
OpenAI unveiled Jalapeño, its first custom inference chip, demonstrating higher throughput per kilowatt and lower latency than commercial alternatives in public benchmarks. The company emphasized an integrated compute strategy spanning hardware, models, and data centers to improve efficiency and cost across workloads.
OpenAI described its compute strategy as a unified system where improvements in one layer—such as custom chips, software, or models—reinforce the others. The company highlighted Jalapeño, its first custom inference chip, which delivered higher peak throughput per kilowatt and lower token latency than commercial systems in the InferenceX benchmark using GPT‑OSS 120B. Performance gains were also observed on DeepSeek R1 and Kimi K2, indicating broader applicability across model families. The chip’s development reflects OpenAI’s focus on integrating hardware, software, and infrastructure to optimize throughput, latency, energy efficiency, and cost for AI workloads.
OpenAI outlined its approach to balancing performance and economics by leveraging a diversified portfolio of hardware and cloud providers, including Microsoft, NVIDIA, AWS, AMD, and others. The company aims to stay on the Pareto frontier by selecting the strongest mix of capability, speed, reliability, efficiency, and cost for each workload. OpenAI actively manages this portfolio to direct demand toward the best performance per dollar while maintaining pricing discipline and adapting to evolving technology. Direct control over components like chips and data centers allows for tighter integration and system-wide improvements.
The company also highlighted Project Camellia, a data center in Georgia designed around customer workloads, featuring job creation, local business support, energy cost coverage, and water conservation through a closed-loop system. The facility’s commitments are subject to annual independent public audits. OpenAI emphasized that the value of this system is measured by its output: more useful intelligence per unit of compute, achieved through better models, smarter routing, optimized software, and purpose-built hardware.
OpenAI noted that improvements in efficiency and capability translate to tangible benefits for customers, such as faster results, fewer retries, and lower total costs. The company cited a 54% reduction in output tokens for GPT‑5.6 Sol with max reasoning on the Artificial Analysis Coding Agent Index compared to another leading model. OpenAI framed these gains as part of a broader economic shift, where greater efficiency expands practical applications, fuels new revenue streams, and funds further investment in research, infrastructure, and safety.