Company

GitLab

Every GitLab engineering case study on TechLogStack — real production incidents, post-mortems, and fixes.

GitLab Deleted Its Own Production Database. Then Found All Five Backups Had Failed Too

On January 31st, 2017, a GitLab engineer trying to fix a lagging database replica ran one command on the wrong host. In two seconds, around 300GB of production data was gone -- and so was the database GitLab.com depended on to function. When the team turned to their backups to recover, they found that of five separate backup and replication mechanisms, only one partial, six-hour-old snapshot actually worked.

production data removed: ~300GB backup mechanisms that had silently failed: 4 of 5 hours to fully restore service: ~18hrs +1 peak viewers on the live recovery stream: ~5,000