DeepSeek overtakes Google on volume, cost per token falls 13.6%
DeepSeek surpassed Google in token volume in July, with its V4 Flash model leading at nearly a fifth of total tokens, while the average cost per token fell 13.6%. Open-weight models like Kimi K3 and GLM 5.2 captured significant workloads, driving volume growth and price declines.
DeepSeek became the second-largest lab by token volume in July, processing more than twice Google’s share, according to Vercel’s AI Gateway data. The company’s V4 Flash model alone accounted for nearly a fifth of all tokens routed through the gateway, surpassing every other model. Google’s token volume share dropped from 24.0% in April to 10.7% in July, while DeepSeek’s rose from less than 1% to 25%. The shift was concentrated in consumer-facing applications, particularly personal-assistant tasks, where DeepSeek’s volume more than tripled.
The average cost per token fell 13.6% in July after rising nearly 20% in May, despite a 37% increase in total spend. Token consumption grew so rapidly that even higher spending did not offset the price drop. Open-weight models, previously dominant only in low-cost segments, now account for 36% of token volume, up from 11% in April. Their share of spending also rose to nearly nine cents of every gateway dollar, the highest in the index’s history, driven by new models like Moonshot’s Kimi K3 and Z.ai’s GLM 5.2.
Anthropic retained the largest share of gateway spending at 29.8%, despite a two-point decline, while OpenAI’s share grew to 12.8%. The four largest frontier labs’ combined spending share fell to 89%, down from a seven-month low of 93%. Google accounted for most of the decline, though DeepSeek’s spending share remained nearly unchanged despite its volume surge. Buyers increasingly routed inference to cheaper models, with OpenAI’s average token cost dropping to 58% of June levels due to a shift toward its cheapest GPT-5-Nano model.
Three-quarters of teams processing over 10 million tokens in both June and July adjusted at least 10% of their model mix, with 60% changing a quarter or more. The median team’s cost per token fell 2.9%, but only one in six teams stayed within 5% of their starting costs. A quarter of teams reduced costs by more than 30%, typically those paying above-average prices, while another quarter paid at least 20% more, often teams that had been below the average. Overall, gateway spend rose 74% since May, and volume doubled, with July’s growth nearly twice June’s pace.