From tokens to tasks: Why agentic AI changes the infrastructure conversation
Arm argues that agentic AI shifts infrastructure focus from raw model speed to end-to-end workflow completion, requiring coordinated CPU, GPU and system-level orchestration to deliver reliable, measurable outcomes.
The useful question is what changes for users, developers or buyers, and whether the announcement stays industry context or becomes something people can actually use.
Agentic AI systems transform isolated prompts into continuous workflows where a single instruction can trigger dozens of coordinated steps—parsing, policy checks, retrieval, model calls, tool validation, sandboxing, edits, tests and verification—making the unit of performance a completed task rather than tokens per second.
Traditional AI infrastructure metrics like tokens per second and latency to first token remain important, but agentic AI expands the critical path to include orchestration, memory, retrieval, tool calls, runtimes, sandboxes, policy enforcement and observability, much of which runs on CPUs.
For users, outcomes matter more than model speed: a developer cares whether a bug was fixed, tests passed and permissions were respected, not raw token throughput, prompting a shift to workflow-level metrics such as cost per completed task, tool-call latency and sandbox startup time.
Arm positions its Neoverse IP, Compute Subsystems and AGI CPU as infrastructure that coordinates heterogeneous systems—CPUs, GPUs, accelerators, memory and networking—to optimize agentic workflows, aiming for measurable business value like lower latency, higher utilization and predictable execution.