How Ora benchmarks every major AI agent on Vercel
Ora, a platform for testing AI agents on live websites, now benchmarks major frameworks including Vercel’s eve on Vercel’s infrastructure. The tool measures success rates, costs, and failures to help companies improve agent readiness.
Ora evaluates how AI agents perform on customer websites by simulating tasks such as signing up, integrating, and paying for products. The platform, launched today, runs agents on live sites via journey.ora.ai, tracking metrics like cost, latency, and steps taken. Ora’s lineup includes frameworks like Claude Code, ChatGPT, and Vercel’s eve, each tested under identical conditions to ensure consistent benchmarking. The results help companies identify where agents fail and what infrastructure changes are needed.
Ora’s testing system is built entirely on Vercel, integrating front-end, back-end, and agent runtime into a single deployment path. Each agent framework requires a separate runtime due to differing tooling and step exposures, but all run on Vercel’s shared infrastructure. Ido Finder, Ora’s engineering lead, highlights the value of side-by-side comparisons, which reveal detailed failure points rather than just scores. The platform’s traces show exactly which steps agents struggle with, enabling targeted fixes to improve performance.
Vercel’s eve framework underwent Ora’s benchmark alongside other agents, using models like Claude Fable 5 and Haiku 4.5. The tests revealed a prompt-caching issue in eve, which the team addressed, reducing total costs by roughly 15% in subsequent runs. Ora’s collaboration with Vercel Engineering allowed direct access to results, accelerating improvements. The benchmark also demonstrated that tasks completed 2x more successfully on customer sites with eve compared to falling back to web search.
Ora is expanding its platform by splitting services into microservices, all deployed on Vercel. The architecture supports internal agents built on eve, which now run as an additional service. Ora estimates that 99% of the web remains unprepared for agent interactions, and its tools aim to bridge that gap. The company’s infrastructure, shared with Vercel, reduces operational overhead, with engineering teams leveraging coding agents to manage deployments efficiently.