OFICIAL Vercel Blog

GLM 5.3 FlashX now available on AI Gateway

What happened
Based on Vercel Blog · Sep 18, 2026

Vercel’s AI Gateway now supports GLM 5.3 FlashX, a high-speed serving option for Z.ai’s multimodal coding model, enabling faster streamed responses for coding agents and interactive applications.

GLM 5.3 FlashX now available on AI Gateway
Vercel Blog — Vercel
Key points
·
GLM 5.3 FlashX delivers inference at approximately 200 tokens per second for faster streamed responses.
·
Users must run vercel ai-gateway setup and select zai/glm-5.3-flashx to configure the model in coding agents.
·
AI Gateway does not charge a platform fee on inference, including Bring Your Own Key requests.
Key numbers
·
The model delivers inference at approximately 200 tokens per second, significantly improving response times for streamed outputs.
·
3 FlashX in a coding agent, users must follow the coding agents guide and execute the command vercel ai-gateway setup to generate a key and configure supported agents.
·
The model is selected within the agent by specifying zai/glm-5.

Vercel has integrated GLM 5.3 FlashX into its AI Gateway, offering a high-speed serving option for Z.ai’s multimodal coding model. The model delivers inference at approximately 200 tokens per second, significantly improving response times for streamed outputs. This enhancement is particularly beneficial for coding agents, tool loops, and interactive applications where users rely on rapid generated output. Developers can access the model through the AI Gateway’s unified API, which simplifies integration and configuration.

To implement GLM 5.3 FlashX in a coding agent, users must follow the coding agents guide and execute the command vercel ai-gateway setup to generate a key and configure supported agents. The model is selected within the agent by specifying zai/glm-5.3-flashx. This process ensures seamless integration with existing workflows while leveraging the model’s improved performance.

AI Gateway provides a unified API for calling models, tracking usage and costs, and configuring retries, failover, and performance optimizations to enhance uptime beyond provider guarantees. The platform includes built-in custom reporting, API key budgets, routing rules, and other advanced features. This infrastructure supports developers in managing model interactions efficiently while maintaining control over performance and expenses.

GLM 5.3 FlashX is available for testing in the AI Gateway model playground, where users can evaluate its capabilities. The platform also lists all supported language models, allowing developers to compare options. AI Gateway reflects provider pricing without markup and does not charge a platform fee on inference, including for Bring Your Own Key requests, ensuring cost transparency and accessibility.

Original source → Deals on Clipraptor.com →