OpenAI Ultrafast mode now available on AI Gateway
Vercel’s AI Gateway now supports OpenAI’s Ultrafast tier for GPT 6 Astra and GPT 6.1 Sol, enabling faster responses for interactive and coding workflows with tiered billing.
AI Gateway has introduced support for OpenAI’s Ultrafast service tier, designed to accelerate output for interactive applications and rapid coding iterations. The feature is available for GPT 6 Astra and GPT 6.1 Sol models, allowing developers to request faster processing when needed. Ultrafast is accessed through the AI SDK, Chat Completions API, or Responses API, with the latter recommended for workflows involving frequent tool calls to minimize overhead.
For persistent connections, Ultrafast supports WebSocket implementations, including examples using the AI SDK over WebSocket. The tier is available for processing in the US and globally, though requests pinned to unsupported regions such as the EU will default to standard tier processing. Standard processing remains the default when no specific tier is requested, ensuring backward compatibility for existing workflows.
Requests served at the Ultrafast tier are billed at six times the standard per-token rate, reflecting the premium for faster output. If a request falls back to a lower tier due to region or other constraints, it is billed at the rate applicable to the tier actually served. This tiered billing structure allows developers to optimize costs based on performance needs and regional availability.
Developers can review current rates for GPT 6 Astra and GPT 6.1 Sol on their respective model pages, which also provide additional details about the GPT-6 model family. The integration with AI Gateway simplifies access to Ultrafast, enabling faster development cycles without requiring direct OpenAI API management.