The Ronan Point tower collapsed in 1968 because a small gas explosion on the 18th floor blew out load-bearing walls, and the entire corner of the building came down. Sam Newman uses that story to open a conversation about progressive collapse in distributed systems, and the parallel is uncomfortable in the best way. We tend to treat software outages as isolated incidents, but the mechanics are identical: a single component fails, the load shifts to its neighbors, and before anyone can act, the whole structure is gone. Newman is not describing a new problem. He is naming a pattern that has been hiding in plain sight, and for anyone responsible for keeping a platform alive, that framing is the real takeaway.
The practical question Newman pushes us to ask is not "how do we prevent every failure?" but "what happens when a failure occurs?" In civil engineering, the answer comes down to redundancy, ductility, and the ability to absorb local damage without global collapse. In software, that translates to strengthening critical components, isolating failure domains, and reducing the number of interconnections that allow a problem to travel. The AWS outages he references are a perfect case study. They are rarely caused by a single catastrophic bug. They are caused by a small issue that cascades because too many services are tightly coupled, because blast radii are too wide, and because the system has no way to degrade gracefully. Newman's argument is that resilience is not about building stronger parts in isolation. It is about designing the whole so that the failure of one part does not become the failure of everything.
Our honest take is that most engineering teams are still optimizing for the wrong metric. They focus on improving uptime of individual services, but they ignore the topology that connects them. You can have a 99.99% reliable database and still take down your entire product if an upstream service has no timeout, no circuit breaker, and no fallback. Newman's work points to a more useful approach: treat interconnections as part of the attack surface, and design for the worst-case load shift, not the happy path. That means deliberately testing what happens when a dependency goes dark, chaos-engineering your own architecture, and being willing to say no to a new integration if it creates a coupling you cannot afford. It is less glamorous than building new features, but it is the work that keeps you from becoming a post-mortem case study.
What we would tell a reader who asks about this presentation is simple: do not wait for the outage to teach you the lesson. Start by mapping your dependency graph and asking which three nodes, if they failed simultaneously, would take down the system. Then ask what you have done to make sure that cannot happen. Newman's examples from civil engineering are not metaphors for decoration. They are warnings that the cost of ignoring progressive collapse is measured in hours of downtime, lost trust, and the painful realization that you did not need a bigger hammer. You needed a better blueprint. The specific thing to watch for in your own architecture is the single point of failure that everyone has stopped noticing because it has been there so long. That is the wall you should be worried about.
