Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed
OpenAI introduces Ultrafast, a new service tier for GPT‑5.6 Sol in the API, delivering up to 14× faster processing with up to 750 output tokens per second, powered by Cerebras.
OpenAI has launched Ultrafast, a new service tier for its GPT‑5.6 Sol model in the API, offering up to 14 times faster processing than the Standard tier. The Ultrafast tier, powered by Cerebras hardware, generates up to 750 output tokens per second, enabling real-time applications where speed is critical. This marks a shift from the trade-off between speed and model intelligence, allowing businesses to integrate advanced AI into time-sensitive workflows without sacrificing capability. Access is currently limited to a preview group, with broader availability planned as capacity increases.
The new Ultrafast tier is designed to support workflows that require rapid responses, such as incident response, coding, financial research, and customer support. During the preview, OpenAI is collaborating with an initial set of customers to evaluate where the speed improvements deliver the most value. Early testing includes building a 3D warehouse simulator from a single text prompt, demonstrating the model’s ability to handle complex, interactive tasks efficiently. The findings will inform future product development and deployment strategies as OpenAI scales capacity.
OpenAI’s internal teams are also testing Ultrafast mode to assess its impact on productivity in real-world scenarios. For incident response, engineers use the tier to analyze logs, traces, and conversations in real time, reducing the time between identifying an issue and taking action. In research workflows, teams leverage Ultrafast to accelerate data queries, knowledge searches, and information synthesis, enabling faster iteration cycles. The goal is to determine how real-time intelligence can enhance decision-making while maintaining human oversight in critical processes.
The Ultrafast tier represents the next phase of OpenAI’s partnership with Cerebras, focusing on ultra-low-latency inference for its most advanced models. Cerebras is now supporting GPT‑5.6 Sol in Ultrafast mode, delivering the high-speed output necessary for businesses to build responsive products and make faster decisions. While access remains limited during the preview, OpenAI plans to expand availability as infrastructure scales. Customers interested in early access can sign up for updates on the company’s website.