1 min readfrom InfoQ

Presentation: When Incidents Refuse to End

Our take

Marathon incidents—those that persist beyond initial response—expose a critical gap between how work is envisioned and how it’s actually performed. Vanessa Huerta Granda’s presentation, "When Incidents Refuse to End," draws on real-world outages to reveal organizational fragility, human limitations, and the intricate web of system dependencies. Learn how structured endurance, humane team rotations, and robust cross-functional coordination are essential for effective incident response. For a forward-looking perspective on outage prediction, explore our article on Empirik’s $21M launch.
Presentation: When Incidents Refuse to End

Vanessa Huerta Granda's exploration of "marathon incidents" offers a crucial perspective on the realities of modern operational resilience. It’s easy to envision incident response as a series of swift, decisive actions – a clean, logical process. However, Granda’s work, and the increasing frequency of prolonged outages highlighted by recent events, demonstrates that the gap between that idealized process and the lived experience of incident response is often vast. The complexity of modern systems, coupled with the inherent limitations of human teams, creates conditions where incidents can drag on for days, even weeks, exposing vulnerabilities across organizations. This echoes the ambition of startups like Empirik, who are attempting to [Sequoia-incubated Empirik launches with $21M to predict outages before they happen], aiming to anticipate these very issues before they escalate. Similarly, the ongoing threat landscape, as evidenced by the recent arrests in the TeamPCP hacks targeting Mercor and OpenAI [Australian police arrest two over TeamPCP hacks targeting Mercor, OpenAI, and others], underscores the persistent need for robust and adaptable incident response strategies.

The core of Granda’s argument – the need for structured endurance, humane rotations, and cross-functional coordination – isn't a novel concept in theory, but it's often neglected in practice. Too often, incident response is treated as a heroic endeavor, demanding unsustainable levels of effort from individuals. This leads to burnout, poor decision-making, and ultimately, prolonged outages. The reality is that managing complex incidents requires a systemic approach, one that acknowledges the cognitive load on responders and prioritizes their well-being. Furthermore, the interconnectedness of modern systems means that incidents rarely remain isolated; they ripple through organizations, impacting teams and processes in unexpected ways. This highlights the importance of breaking down silos and fostering a culture of shared responsibility during times of crisis. The impact of these events can be significant, as seen with Boston Scientific’s recent cyberattack [Medical device maker Boston Scientific says a cyberattack is causing a ‘global disruption’ to its operations], demonstrating the potential for widespread operational and reputational damage.

What's particularly insightful about Granda’s framework is its focus on the “work as done” versus “work as imagined” discrepancy. Incident response plans are often created in a vacuum, assuming a level of predictability and control that simply doesn’t exist in the real world. Marathon incidents expose the fragility of these plans, revealing the limitations of our understanding of system interdependencies and human behavior under pressure. This isn’t about fault or blame; it’s about recognizing that complex systems are inherently unpredictable and that effective incident response requires adaptability, continuous learning, and a willingness to challenge assumptions. The move towards AI-powered monitoring and prediction tools, while promising, must be coupled with a deeper understanding of the human element in incident response—the fatigue, the biases, the communication breakdowns that can exacerbate already challenging situations.

Ultimately, Granda’s work calls for a shift in how we approach operational resilience. It’s not just about building more robust systems; it’s about building more resilient organizations—ones that can withstand prolonged disruptions, learn from their mistakes, and support the people who are on the front lines of incident response. As systems become increasingly complex and interconnected, and the threat landscape continues to evolve, the ability to effectively manage marathon incidents will be a defining factor in the success of businesses across all sectors. The question now is: how can organizations move beyond reactive measures and proactively cultivate the structures, processes, and culture needed to truly thrive in an era of persistent operational challenges?

Vanessa Huerta Granda explains how marathon incidents expose the gap between work as imagined and work as done. Drawing from real-world scenarios, she shares how complex outages reveal organizational fragility, human limits, and system interdependencies—and why incident response requires structured endurance, humane rotations, and holistic cross-functional coordination.

By Vanessa Huerta Granda

Read on the original site

Open the publisher's page for the full experience

View original article