Most engineering leaders treat incident counts the way drivers treat a flashing oil light: as a warning that something is breaking. So when Great Circle published its argument that rising incident numbers can signal an improving reliability culture, it likely made some engineering VPs uncomfortable. That discomfort is worth sitting with. If you are leading a platform team and the numbers go up, your first instinct is to brace for a difficult quarterly review. The more useful instinct, as Great Circle suggests, is to ask what those numbers represent. Are more failures happening, or are more failures being found? Those are not the same question, and confusing them has caused more than a few misguided reliability initiatives.
This connects to a broader theme we have been tracking across our coverage. In Refine Your Accepted Paper: Maximizing Changes Before Camera Ready, we explored how late-stage revisions can feel like failure when they actually represent a healthy engagement with feedback. The same logic applies to incident reporting. A team that hides a production issue because it fears consequences is not more reliable; it is merely quieter. A team that surfaces issues early and often is building the kind of institutional memory that prevents recurrence. Similarly, Beyond Green: Understanding True Test Suite Efficacy challenged the assumption that a passing test suite means quality. Green tests can be lying to you. Incident counts can do the same, but in the opposite direction. They can make a healthy system look unstable when the real problem is that nobody was counting before.
Our honest take is that Great Circle is pointing at a maturity model, not a bug. The organizations that see rising incident counts as a problem are often the same ones that have not built a blameless postmortem culture. They have not created the psychological safety required for engineers to say, "I broke this, and here is what I learned." When that culture exists, incident reports become a form of documentation, not a confession. The metric is a proxy for trust. So what should a leader actually do when incidents rise? Resist the urge to punish the reporting mechanism. Instead, look at the severity distribution. If the increase is concentrated in low-severity, low-blast-radius events, you are likely seeing the system surface what it used to swallow. That is progress, even if it feels like regression.
The practical takeaway here is direct: stop managing to the number and start managing to the response. A rising incident count is only meaningful relative to your team's ability to detect, respond, and learn. If you have to choose between a team that reports everything and a team that reports nothing, the former is healthier even when it looks messier. The real question to watch is whether the incidents you are seeing now are the same ones you saw six months ago. If they are, your culture is improving but your engineering is not. That is the detail worth tracking.
