OFICIAL GitHub Blog

The August 17 outage, and the work ahead

What happened
Based on GitHub Blog · Aug 20, 2026

GitHub reported a 7-hour, 47-minute outage on August 17 affecting core services including authentication, Actions, APIs, and Copilot, following a similar incident on August 6. The company attributed both failures to capacity constraints in its Central US data center during peak traffic.

The August 17 outage, and the work ahead
GitHub Blog — GitHub
Key points
·
An update on the August 17 outage and the steps the company is taking to improve reliability.
·
On August 17, GitHub experienced an outage that lasted 7 hours and 47 minutes.
·
It disrupted github.com, authentication, GitHub Actions, APIs, pull requests, issues, and Copilot, affecting developers and organizations around the world.
·
If you the company is trying to ship software that day, we let you down.
Key numbers
·
com, authentication, GitHub Actions, APIs, pull requests, issues, and Copilot for 7 hours and 47 minutes, impacting developers globally.
·
4 billion to 2.
·
9 billion, straining infrastructure despite ongoing reliability efforts.

GitHub confirmed that the August 17 outage disrupted github.com, authentication, GitHub Actions, APIs, pull requests, issues, and Copilot for 7 hours and 47 minutes, impacting developers globally. The incident followed an August 6 failure in Actions and marked the second major disruption in August. GitHub acknowledged that the outage began when traffic reached a new peak, overwhelming a critical infrastructure component in its Central US data center, which failed to scale appropriately.

An investigation revealed that the capacity pressure cascaded through systems, causing authentication failures and service disruptions. Recovery required rerouting traffic, isolating affected infrastructure, and staged service restorations, with Copilot services taking longer due to client-side retry loops that increased traffic during recovery. The root cause analysis identified neither code nor configuration changes as the cause, but rather a failure to scale critical components ahead of demand.

Since April, monthly commits on GitHub have surged from 1.4 billion to 2.9 billion, straining infrastructure despite ongoing reliability efforts. GitHub has added over 3 million CPU cores, 120 petabytes of high-speed storage, and expanded network capacity, while accelerating migration to Azure, which now handles roughly 58% of platform load and half of all Git operations. The company is also addressing operational gaps by improving testing, rollouts, observability, and alerting, while isolating critical systems to reduce outage likelihood and impact.

Immediate changes include implementing consistent retry limits and budgets across service interactions to prevent retry storms, and reviewing lower-priority alerts to identify components vulnerable to sudden traffic spikes. GitHub emphasized that reliability is a continuous commitment, acknowledging the August 17 incident as a failure to meet community expectations. The company plans to roll out an architecture enabling linear scaling of read capacity, beginning with the largest monorepos.

Original source → Deals on Clipraptor.com →