workflow automation

From reactive fixes to self-healing data: inside Netflix's observability knowledge graph

Netflix processes 38 million events per second, and its engineers were tired of playing catch-up.

3 min readInfoQ
From reactive fixes to self-healing data: inside Netflix's observability knowledge graph

Netflix processes 38 million events per second. That number alone should make every engineering leader pause, because it represents the difference between observability as a cost center and observability as a competitive advantage. The company's move to unify metrics, events, logs, and traces into a queryable knowledge graph, then layer agentic workflows on top, is not just a technical upgrade. It is a philosophical shift from asking "what broke?" to asking "what will break, and how do we prevent it?"

The practical implication for teams outside Netflix is immediate and uncomfortable: reactive monitoring is a legacy tax. When you wait for alerts to fire, you are already behind. Netflix's approach, detailed by Prasanna Vijayanathan and Renzo Sanchez-Silva, treats telemetry not as isolated signals but as a connected web of operational truth. By mapping relationships between services, dependencies, and historical incidents, their AI-driven ontology can triage issues and suggest root causes before a human even opens a dashboard. The goal is self-healing systems, where the infrastructure itself becomes the first responder.

This is where the comparison to Anthropic updates rules targeting election interference and model abuse becomes instructive. Anthropic is defining the boundaries of what Claude can and cannot do in high-stakes contexts. Netflix is doing the opposite: expanding Claude's autonomy inside its own operational environment. Both are wrestling with the same underlying question, which is how much agency we grant AI systems. But the stakes differ. An election interference policy protects democratic processes; a self-healing observability graph protects uptime. One is about constraint, the other about enablement, and both are necessary for AI to be trusted with more responsibility.

What makes this work feel genuinely progressive, rather than another vendor pitch, is the emphasis on accessibility. The engineers are not describing a black box that magically fixes problems. They are building a knowledge graph that humans can query, interrogate, and refine. That is a crucial distinction. The tool does not replace the operator; it augments the operator's ability to reason about complex systems. This aligns with a broader cultural moment where we are reassessing what intelligence actually looks like in practice. Even Ben Affleck brings unexpected credibility to thoughtful AI conversations by pointing out that AI is not magic, but a tool that requires understanding to wield effectively. Netflix's engineers are demonstrating that principle in production.

The takeaway is specific: start mapping your telemetry relationships now, before the volume of events outpaces your team's ability to reason about them. You do not need Netflix-scale infrastructure to benefit from a knowledge graph approach. You need the discipline to connect your data points, rather than storing them in silos. The open question is whether the broader industry will follow this lead or continue paying the reactive tax. Watch how your own monitoring tools evolve over the next year. If they are not helping you ask better questions, they are just generating more noise.

From InfoQ

Prasanna Vijayanathan and Renzo Sanchez-Silva share how Netflix tackles observability across 38M events/sec. They discuss replacing reactive monitoring with an AI-driven operational ontology and agentic workflows using Claude and graph databases. They explain how unifying MELT telemetry into queryable knowledge graphs enables automated triaging, root-cause analysis, and self-healing systems.

Read the original at InfoQ