AI

How Expedia's AI Observability Platform Helps Engineers Investigate Faster

When a production incident hits, every minute of investigation matters.

3 min readInfoQ
How Expedia's AI Observability Platform Helps Engineers Investigate Faster

Expedia Group's new STAR platform is a practical answer to a problem every engineering team knows too well: the chaotic, draining hunt for root cause during a production incident. STAR, built on FastAPI, Datadog, Celery, Redis, and Langfuse, applies structured workflows to service telemetry, letting large language models surface root cause assessments while keeping a human engineer firmly in the loop. That last part matters. Too often, AI tools are positioned as replacements for judgment, but STAR treats the engineer as the decision-maker and the model as a high-speed analyst. It's a sensible division of labor, and it's one that reflects a mature understanding of how these systems actually perform under pressure.

The timing is interesting, especially given how we've been circling the question of trust in AI-assisted workflows. A related piece on Talking to My AI Clone Taught Me to Question the Tech highlights the unease that comes when AI output feels plausible but isn't verifiable. STAR sidesteps that trap by design. It doesn't just hand an engineer a conclusion; it structures the analysis so a human can trace the reasoning, challenge it, and override it. That's the difference between using AI as a crutch and using it as a lever. For teams watching this space, the lesson isn't about the specific stack or the LLM choice. It's about the workflow design: keeping the human accountable while letting the machine do the heavy lifting on data correlation.

What we find genuinely useful here is the emphasis on structured workflow over raw model power. Anyone who has worked with LLMs knows they're impressive at pattern matching but unreliable on their own. By embedding the model inside a defined incident response process, Expedia is treating AI as a component of a larger system, not as a magic wand. That's a pragmatic stance. It also echoes a theme from our practical guide to distributed algorithms, where the focus is on understanding the underlying mechanics rather than just trusting the output. STAR works because it forces the model to operate within boundaries that produce auditable, explainable results. That's not flashy, but it's effective.

The open question we'd raise is about generalization. STAR is built for Expedia's environment, with its specific services, its telemetry, and its incident culture. Can that approach be abstracted into a product other teams can adopt? Or is the value mostly in the discipline of building it in-house? The answer will determine whether STAR is a one-off success or a template for the industry. For now, the concrete takeaway is this: the next time you're tempted to drop an LLM into your incident response, ask not what the model can tell you, but how you'll verify what it says. That's the standard STAR sets, and it's the right one. We'll be watching to see if the pattern spreads beyond Expedia's walls.

From InfoQ

Expedia Group has introduced STAR, an internal AI-assisted observability platform that helps engineers investigate production incidents using service telemetry and LLMs. Built with FastAPI, Datadog, Celery, Redis, and Langfuse, STAR follows structured workflows to analyze telemetry, generate root cause assessments, and support incident response while keeping engineers in the loop.

Read the original at InfoQ