Hacker News
The August 17 outage, and the work ahead
On August 17, GitHub experienced a 7-hour-47-minute outage after a critical component in its Central US data center failed to scale under record traffic, causing authentication failures that disrupted GitHub Actions, APIs, pull requests, issues and Copilot. The outage, like the August 6 incident, resulted from capacity shortages despite unchanged code or configuration, leading to immediate actions: enforcing consistent retry limits, tightening alert reviews, and accelerating capacity additions—adding over 3 million CPU cores, 120 PB of high-speed storage, and shifting roughly 58% of the platform’s load to Azure.