High Availability

When health checks lie, high availability becomes a hollow promise

Most system outages aren't caused by hardware failure.

3 min readInfoQ
When health checks lie, high availability becomes a hollow promise

A routine TLS 1.3 upgrade should not be the event that takes down your production traffic, yet that is exactly what Alexey Golev documents in his recent analysis. The failure mode is instructive and unsettling: Route 53 health checks broke silently during the upgrade, a CDN stopped routing traffic to a perfectly healthy region, and internal dashboards reported no anomalies. This is not a story about a bad deployment. It is a story about how we confuse high availability with resilience, and how that confusion creates invisible failure modes that no dashboard can catch.

The distinction Golev draws matters deeply for any team building modern infrastructure. High availability assumes your components are working and simply need redundancy. Resilience assumes components will fail in unpredictable ways and that the system must recover. The TLS 1.3 incident is a textbook case of the gap between these two concepts. The health checks were available, but they were lying. The control plane appeared operational, but its dependencies had shifted under the surface. This is the kind of failure that erodes trust in your monitoring faster than a full outage ever could. For teams exploring the Explore the Future of AI Deployment: Key Topics at QCon AI New York, the lesson is that production guardrails must account for control-plane dependencies, not just data-plane throughput.

The practical consequence for our readers is direct. If your team owns a service with health checks, traffic routing, or any automated failover logic, you now have a concrete scenario to test. Run a TLS upgrade in a staging environment and watch what happens to your health check endpoints. Do your dashboards report success when the underlying check is actually returning stale or malformed data? If so, you have discovered a resilience gap, not an availability problem. This is where the Scale Sandboxes Instantly: A New Approach to Concurrent AI Workloads article offers a complementary insight: sandbox environments are precisely the right place to inject these kinds of control-plane failures without risking production. Treat your health check infrastructure as a workload that needs its own isolation and testing.

Golev also highlights a subtler erosion: recovery capability disappears when no one explicitly owns it. After the incident, the team likely fixed the health check, but who is responsible for ensuring the next silent failure is detected before it causes harm? Without explicit ownership of recovery processes, teams default to patching the symptom and moving on. The one concrete takeaway here is that every health check should have a documented failure mode that is exercised at least quarterly, and that exercise should include verifying the monitoring layer itself. A dashboard that shows green while traffic is being dropped is not a dashboard at all, it is a liability.

From InfoQ

A routine TLS 1.3 upgrade silently broke Route 53 health checks, causing a CDN to stop routing traffic to a healthy region while internal dashboards showed nothing wrong. This article examines why HA and resilience are different problems, how control-plane dependencies create invisible failure modes, and why recovery capability erodes without explicit ownership.

Read the original at InfoQ