1 min readfrom InfoQ

More Incidents Don't Necessarily Mean Less Reliability

Our take

A common misconception in engineering leadership is that more reported incidents equate to lower system reliability. Recent analysis, however, suggests the opposite: a rising incident count often reflects an *improving* incident management culture—organizations are better at identifying and reporting issues. This indicates greater visibility and proactive problem-solving. Explore this counterintuitive insight further, and consider how embracing robust incident reporting can ultimately strengthen your systems. For a deeper dive into related technological shifts, see our article on Netflix's adoption of Kueue.
More Incidents Don't Necessarily Mean Less Reliability

The conventional wisdom in engineering leadership – that more incidents equate to less reliable systems – is a surprisingly persistent one. It’s a gut reaction, fueled by the desire to project stability and confidence. However, as highlighted in a recent piece from Great Circle, this assumption can be fundamentally misleading. In fact, a rising tide of reported incidents might actually signify a maturing and increasingly robust incident management culture. This shift is particularly relevant now, as organizations grapple with the complexities of AI-native infrastructure and increasingly distributed systems – complexities that are only amplified by the innovative approaches described in Writer introduces new AI model and upgraded harness to contain token costs. The key isn’t simply minimizing incident counts, but fostering an environment where issues are openly reported, thoroughly investigated, and ultimately, used to improve overall system resilience.

The core of the argument rests on the idea that a culture of silence around incidents – often driven by fear of blame or career repercussions – is far more detrimental to reliability than the incidents themselves. When teams are discouraged from reporting problems, underlying systemic issues remain hidden, festering and potentially leading to larger, more catastrophic failures down the line. Conversely, when reporting is encouraged and viewed as a learning opportunity, organizations can identify and address root causes proactively. Netflix’s move to adopt Kueue, as detailed in Netflix Adopts Cloud-Native Job Queueing System Kueue to Replace an In-House Solution, exemplifies this principle; embracing an open-source solution allowed them to leverage a broader community's experience and accelerate improvements to their batch job execution system, a process likely informed by diligent incident analysis. This shift requires a significant cultural investment, promoting psychological safety and celebrating learning from mistakes – a vital consideration for founders navigating the fast-paced environment described in The founder’s guide to TechCrunch Disrupt 2026: Everything you need to know.

This isn’t to say that incident counts are irrelevant. They remain a valuable data point, but their interpretation needs to be nuanced. A sudden, dramatic spike should always trigger immediate investigation. However, a gradual increase, particularly when accompanied by improvements in incident response time and root cause analysis, should be viewed as a positive sign. The focus should be on the *quality* of incident management, not just the quantity of incidents. This requires investing in tools and processes that facilitate rapid diagnosis, effective communication, and actionable insights. It also necessitates empowering teams to own their systems and take responsibility for their performance. The move away from reactive firefighting toward proactive resilience is a hallmark of mature engineering organizations.

Ultimately, the Great Circle article serves as a potent reminder that reliability isn't about achieving a state of incident-free perfection, a goal that’s increasingly elusive in today’s complex technological landscape. It’s about building systems that are designed to fail gracefully, and cultivating a culture that embraces failure as an opportunity for growth. As we move further into an era defined by AI-powered systems and ever-increasing complexity, the ability to learn from incidents and continuously improve will be the defining factor separating resilient organizations from those that crumble under pressure. The question now is: how can engineering leaders effectively measure and reward the *process* of incident management, rather than solely focusing on the outcome of minimizing incident counts?

One of the most common assumptions in engineering leadership is that a rising number of reported incidents signals declining system reliability. However, a recent article from Great Circle argues that the opposite is often true: an increase in incident counts may actually indicate that an organization's incident management culture is improving.

By Craig Risi

Read on the original site

Open the publisher's page for the full experience

View original article