OFICIAL Microsoft Azure Blog

AT&T and Microsoft scale trillion-token workloads with Microsoft Foundry and AMD

What happened
Based on Microsoft Azure Blog · Jul 23, 2026

AT&T and Microsoft scaled trillion-token AI workloads for telecom-focused models using Microsoft Foundry and AMD GPUs, enabling cost-efficient, flexible development of domain-specific AI systems.

AT&T and Microsoft scale trillion-token workloads with Microsoft Foundry and AMD
Microsoft Azure Blog — Microsoft
Key points
·
Telecommunications organizations are increasingly looking to AI to help teams navigate highly specialized domains, but generic models often lack the industry-specific knowledge needed to understand telecom networks, standards, and operations.
·
To address that gap, AT&T created their Open Telco (OTel) models, the next generation of telecom-focused AI designed to bring deeper telecommunications expertise into AI systems.
·
Building OTel2.0 required more than training a large language model, it reflected a broader issue many organizations face: how to build domain-specific AI systems at scale while balancing cost, performance, and operational complexity.
·
To continue advancing telecom-focused AI, AT&T needed a platform capable of supporting OTel2.0 development at an entirely new scale.
Key numbers
·
AT&T adopted a multi open-model strategy, deploying models like Phi-4, OSS-120B, and Gemma-4 from Hugging Face via Microsoft Foundry.
·
Phi-4 processed over 700 billion tokens monthly for data preparation and training.
·
The project utilized approximately 530 GPUs, including 430 AMD Instinct MI300X GPUs, across heterogeneous architectures to optimize model deployment and performance.

AT&T developed OTel2.0, a telecom-specific AI model, to address gaps in generic AI systems that lack industry expertise. The project required scalable infrastructure to handle massive data volumes while balancing cost and performance. Microsoft Foundry Managed Compute provided dedicated GPU capacity, eliminating the need for AT&T to manage deployments and infrastructure overhead, streamlining development processes.

AT&T adopted a multi open-model strategy, deploying models like Phi-4, OSS-120B, and Gemma-4 from Hugging Face via Microsoft Foundry. Phi-4 processed over 700 billion tokens monthly for data preparation and training. The approach allowed AT&T to tailor workflows for telecom-specific needs, control costs, and accelerate innovation while maintaining flexibility in model selection and deployment.

The project utilized approximately 530 GPUs, including 430 AMD Instinct MI300X GPUs, across heterogeneous architectures to optimize model deployment and performance. Microsoft Foundry enabled rapid deployment in days rather than weeks, supporting AT&T’s need for speed and scalability. This infrastructure flexibility aligns with broader industry trends, where organizations prioritize model choice, cost optimization, and operational efficiency in AI development.

AT&T processed around 1 trillion tokens for OTel2.0, combining raw documents from GSMA with synthetic data generated using open models like Phi-4. The approach saved tens of millions of dollars compared to using frontier models, allowing teams to focus on larger-scale experimentation and business value. Microsoft Foundry’s managed compute platform proved critical in transforming infrastructure from a deployment consideration into a strategic component of AI development.

Original source → Deals on Clipraptor.com →