MAI-Code-1.1-Flash: Better, faster, at a quarter of the cost
Microsoft released MAI-Code-1.1-Flash, a smaller and more efficient AI coding model, reducing costs by 75% while improving performance in GitHub Copilot.
Microsoft announced MAI-Code-1.1-Flash, a new AI coding model that delivers higher-quality code at 25% greater token efficiency and a quarter of the cost compared to its predecessor, MAI-Code-1.0, launched in June. The update prioritizes developer feedback, particularly for CLI tasks and .NET performance, resulting in a 22% improvement on Terminal-Bench 2.1 and a 15% boost on .NET tasks within GitHub Copilot. The model is now available in production for GitHub Copilot users.
The improvements extend beyond benchmarks, with a 4% increase in code survival rates and a 9% rise in return visits, indicating better real-world reliability. MAI-Code-1.1-Flash also processes tokens 25% faster while using 25% fewer tokens per task, reducing latency and computational costs. Microsoft attributes these gains to optimized training and serving efficiency, achieved through reinforcement learning across hundreds of thousands of environments in GitHub Copilot.
Pricing for MAI-Code-1.1-Flash is set at one quarter of the cost of MAI-Code-1.0, reflecting the model’s efficiency gains. Microsoft emphasized that the update was driven by iterative development, with a focus on practical use cases rather than simply scaling up model size. The company described the process as a continuous loop of shipping, learning, and improving, leveraging a lean team of researchers and engineers.
Microsoft invited users to test MAI-Code-1.1-Flash in GitHub Copilot and provide feedback by opening issues on the project’s repository. The company positioned itself as a talent-dense team aiming to build advanced AI models while maintaining operational efficiency. Microsoft also expressed openness to hiring, highlighting its mission-driven approach to AI development.