Company

AWS

Every AWS engineering case study on TechLogStack — real production incidents, post-mortems, and fixes.

AWS Infrastructure Outages
15 min

Why Did AWS Overheat? Inside the North Virginia Thermal Surge and Power Failures That Halted Global Trading Rails

On May 7, 2026, a sudden cooling failure at an Amazon Web Services (AWS) facility in Northern Virginia turned the backbone of the internet into an environmental oven. High temperatures triggered a critical 'thermal event' and subsequent power loss, causing widespread instance impairments and freezing operations for enterprise heavyweights like Coinbase, FanDuel, and CME Group. Far from a routine glitch, this high-profile disruption exposed the extreme fragility of high-density availability zones failing under synchronous load.

{'label': 'peak incident duration', 'value': '21 Hours'} {'label': 'affected core services', 'value': '12+'} {'label': 'downdetector alert spike', 'value': '~600 Blocks'} +1 {'label': 'availability zones isolated', 'value': '1 Zone'}

One Wrong Number in a Routine Command Took Down Slack, Trello, and a Chunk of the Internet

On the morning of February 28th, 2017, an authorized S3 engineer ran a command to remove a small number of servers from a billing subsystem. One input was typed wrong. The command removed far more capacity than intended -- enough to force two core S3 subsystems into a full restart neither had needed in years. For the next four hours, US-EAST-1 was effectively offline, and a long list of sites that quietly depended on S3 went down with it.

PST time the command ran: 9:37 AM total disruption: 4hr 17min core subsystems needing full restart: 2 +1 high-profile sites affected: 100+