Designing For Failure: 4 Resilience Practices That Make Outages Boring
devops.com, December 10,2025
Last winter, my city Richmond VA suffered water distribution outages for days after a blizzard. Not because of one big failure, but because backup pumps failed, sensors misread, alerts got buried, and then another pump died during recovery.
The whole city ended up under a boil‑water advisory. Sound familiar? Replace 'water pumps' with 'microservices' and you've got every cascading outage I've debugged in over 15 years.
The timeline mapped perfectly to Dr. Richard Cook's observations on complex systems: failures are multi‑factor, systems constantly run in degraded mode (not everything is perfect all the time), and single‑cause incident findings often mislead. Modern outages aren't 'one bad deploy.' They're three or four small issues that line up at exactly the wrong moment. It could be timing or load and a dependency wobble, an alert nobody saw because they were fixing something else.