Asana cuts model costs 76x in browser tests with GPT-6.1 Sol
Asana reduced browser agent costs by 76x and improved speed fivefold by optimizing workflows with OpenAI’s GPT‑6.1 Sol and GPT‑6 Astra in Codex.
Asana integrated OpenAI’s GPT‑6 Astra in Codex to analyze and optimize its browser agent workflows, achieving a 76-fold reduction in model costs and a fivefold increase in speed during tests. The company, which provides automation across business applications via its StackAI platform, sought to address inefficiencies in workflows that navigate websites, fill forms, and gather data without coding. By leveraging GPT‑6 Astra, Asana automated the investigation and testing process, cutting manual effort from one to two months down to about a week.
The optimization study, involving 144 runs, compared GPT‑6.1 Sol with three other frontier models. The revised workflow on GPT‑6.1 Sol averaged $0.47 in model costs and completed tasks in roughly four minutes, compared to the original setup on Model B, which cost at least $36.21 per run and took over 22 minutes. GPT‑6 Astra identified inefficiencies in how the agent managed browsing history and screenshots, leading to a new caching and screenshot policy that reduced costs while preserving accuracy.
GPT‑6 Astra’s experiments tested two history budgets and six caching policies, with the best-performing approach allowing screenshots to accumulate to 20 before trimming. This policy, combined with a larger history budget, became the optimized workflow. Each configuration performed the same task: collecting six fields for 32 books from a public demo catalog, representative of customer workflows in StackAI. All runs in the optimized workflow completed successfully and returned correct answers.
The findings were implemented in production, reducing Model B’s estimated cost from at least $36.21 to $1.24 per run, a 29-fold improvement, while GPT‑6.1 Sol further lowered costs to $0.47. Asana has released the changes to browser navigation in StackAI and is developing tools to replicate such experiments. The team plans to integrate this testing into platform evaluations, enabling comparisons of cost, runtime, and answer quality for agent configurations.