AMD and Cerebras Announce Industry-Leading Ultra-Low-Latency and High Throughput AI Inference Solution
AMD and Cerebras unveiled a disaggregated AI inference solution combining AMD Helios rack-scale systems with Cerebras Wafer-Scale Engine technology, targeting ultra-low-latency and high-throughput workloads.
AMD and Cerebras Systems announced a partnership to deliver a disaggregated AI inference solution that integrates AMD Helios rack-scale systems with Cerebras Wafer-Scale Engine technology. The solution, unveiled at Advancing AI 2026, aims to meet the dual demands of ultra-low latency and high throughput required for advanced AI applications. AMD Helios provides scalable throughput, while the Cerebras WSE handles memory-bandwidth-intensive token generation with minimal delay. Together, the companies claim the combined system can achieve up to 5x higher tokens per second per watt compared to existing configurations.
The joint solution addresses the growing diversity in AI inference workloads, where some applications prioritize token generation volume while others require faster response times. AMD Helios serves as the high-throughput prompt processing engine, capable of handling large context windows and complex requests at scale. Cerebras WSE accelerates the decode and token generation phase, ensuring real-time performance for latency-sensitive tasks such as coding assistants, autonomous agents, and live agent workflows. The disaggregated approach allows each compute engine to be optimized independently for its specific role in the inference pipeline.
AMD and Cerebras plan to deploy the solution initially through Cerebras Cloud in the second half of 2026, with Cerebras integrating AMD Helios systems into its data centers. The collaboration targets the ultra-low-latency segment of the inference market, offering a platform designed to balance high throughput with real-time responsiveness. AMD’s CEO, Dr. Lisa Su, emphasized the need for flexible infrastructure to meet the evolving demands of AI inference, while Cerebras CEO Andrew Feldman highlighted the growing demand for ultra-fast inference across industries.
The announcement follows modeling by AMD Performance Labs and Cerebras in July 2026, which compared the joint solution’s tokens per second per kilowatt performance against a Cerebras WSE-only configuration using the Kimi 2.6 1T Model. The companies note that actual results may vary depending on system configurations. The solution is positioned to support applications in software development, robotics, and scientific discovery, where response time directly impacts user experience and system utility.