OFICIAL Vercel Blog

Qwen 3.8 Flash now available on AI Gateway

What happened
Based on Vercel Blog · Aug 26, 2026

Vercel’s AI Gateway now supports Alibaba’s Qwen 3.8 Flash model, enabling text and image input with a 1 million token context window and 65k token output. The model targets coding, tool use, and multi-step agent workflows.

Qwen 3.8 Flash now available on AI Gateway
Vercel Blog — Vercel
Key points
·
Qwen 3.8 Flash from Alibaba is now available on AI Gateway.
·
It takes text and images as input, serves a context window of 1 million tokens, and can return up to 65k tokens in a response.
·
Alibaba recommends it for coding, tool use, and multi-step agent workflows.
·
To use it in a coding agent, see the coding agents guide, then run vercel ai-gateway coding-agents setup to connect agents like Claude Code, Codex, OpenCode, Cursor, Pi, and more and select alibaba/qwen3.8-flash inside the agent.
Key numbers
·
The model accepts both text and images as input, supports a context window of up to 1 million tokens, and can generate responses of up to 65,000 tokens.
·
8 Flash model, enabling text and image input with a 1 million token context window and 65k token output.

Vercel has integrated Alibaba’s Qwen 3.8 Flash into its AI Gateway, expanding the platform’s model offerings. The model accepts both text and images as input, supports a context window of up to 1 million tokens, and can generate responses of up to 65,000 tokens. Alibaba positions the model for coding tasks, tool integration, and multi-step agent workflows, aligning with developer needs for high-capacity processing.

To deploy Qwen 3.8 Flash in coding agents, developers must follow Vercel’s coding agents guide and execute the command vercel ai-gateway coding-agents setup. This process enables connections to agents such as Claude Code, Codex, OpenCode, Cursor, and Pi, with the model accessible under the identifier alibaba/qwen3.8-flash. The setup streamlines integration while maintaining compatibility with existing agent frameworks.

AI Gateway provides a unified API for model interactions, including usage tracking, cost management, and configurable retries, failover, and performance optimizations to enhance uptime beyond provider guarantees. The service includes built-in custom reporting, Zero Data Retention support, budget controls for API keys, and routing rules to manage traffic and costs efficiently.

Pricing for Qwen 3.8 Flash on AI Gateway reflects the provider’s rates without additional markup, and the platform does not charge a fee for inference, even for Bring Your Own Key (BYOK) requests. This pricing structure aims to reduce costs for developers while ensuring transparent billing aligned with Alibaba’s published rates.

Original source → Deals on Clipraptor.com →