GLM 5.3 now available on AI Gateway
Z.ai’s GLM 5.3 model is now accessible via Vercel’s AI Gateway, offering enhanced performance in software engineering and multi-step agent tasks while reducing output tokens. The update includes improved vulnerability discovery capabilities.
GLM 5.3 from Z.ai has been added to Vercel’s AI Gateway, introducing performance improvements over its predecessor, GLM 5.2, particularly in complex software engineering tasks and multi-step agent workflows. The model achieves these gains while generating fewer output tokens at equivalent effort levels, according to Z.ai. It also demonstrates stronger capabilities in identifying vulnerabilities across exploitation chains, as measured by DeepsecBench, which evaluates security flaw detection in application code and associated computational costs.
The model maintains the same input and output specifications as GLM 5.2, supporting a 1 million token context window and a maximum output of 128,000 tokens. GLM 5.3 retains support for function calling, structured output, streaming, and context caching. Users can integrate it into coding agents such as Claude Code, Codex, OpenCode, Cursor, and Pi by running the Vercel AI Gateway setup command and selecting the zai/glm-5.3 endpoint.
AI Gateway provides a unified interface for model deployment, offering features like usage tracking, cost management, and configuration options for retries, failover, and performance optimizations to enhance uptime beyond provider guarantees. The platform includes built-in custom reporting, Zero Data Retention support, API key budgeting, and routing rules. Pricing aligns directly with provider rates, with no platform markup applied to inference costs, including Bring Your Own Key (BYOK) requests.
AI Gateway’s model leaderboard ranks the most frequently used models based on total token volume processed across all traffic, offering transparency on adoption trends. The addition of GLM 5.3 expands the available model options for developers relying on AI Gateway for scalable and cost-effective AI workloads.