Expedia Uses AI Driven Service Telemetry Analyzer to Accelerate Incident Investigation
Our take

Expedia’s introduction of STAR, their AI-assisted observability platform, signals a significant shift in how large organizations are tackling the complexities of modern incident investigation. The sheer volume of data generated by today’s distributed systems makes manual root cause analysis increasingly unsustainable. STAR’s approach, leveraging LLMs and structured workflows alongside established tools like FastAPI, Datadog, Celery, Redis, and Langfuse, represents a practical evolution, not a radical departure. It’s a move that aligns with the growing recognition that AI isn't meant to replace engineers, but to empower them. The timing is particularly relevant considering the discussions happening around production AI, as highlighted in QCon AI New York 2026: Registration Opens for December 15-16 Production-AI Conference, indicating a broader industry focus on operationalizing AI within critical systems. Furthermore, the efficiency gains STAR promises resonate strongly with the ongoing search for improved developer productivity, a topic explored in 7 Best Claude Code Alternatives for CLI Agentic Coding, where developers are actively seeking tools to streamline their workflows.
The key to STAR’s potential lies in its emphasis on "keeping engineers in the loop." Observability platforms that completely automate incident resolution risk introducing new vulnerabilities and eroding trust. STAR's structured workflows and AI-generated root cause assessments serve as intelligent assistants, providing valuable context and accelerating investigation, but ultimately leaving the final decision-making authority with human engineers. This is crucial for maintaining system integrity and ensuring that AI-driven insights are properly vetted. The integration with existing telemetry tools like Datadog also demonstrates a pragmatic approach, avoiding the need for wholesale platform replacements, which are often met with resistance due to cost and disruption. The fact that Expedia built this internally also speaks volumes – it suggests that off-the-shelf solutions weren't adequately addressing their specific needs, a common challenge for organizations of their scale and complexity.
The broader implications of STAR's development extend beyond Expedia. We’re seeing a growing trend of organizations building custom AI-powered observability tools, often leveraging open-source LLMs and frameworks to tailor solutions to their unique environments. This reflects a maturing understanding of AI's role in operations – it's not about replacing existing infrastructure, but about augmenting it with intelligent layers that can automate repetitive tasks, surface critical insights, and ultimately improve system resilience. The choice of technologies like FastAPI for the backend highlights a focus on performance and scalability, essential for handling the high throughput of telemetry data. The careful selection of components suggests a considered engineering approach, prioritizing reliability and maintainability alongside AI capabilities.
Looking ahead, the success of platforms like STAR will depend on their ability to continuously learn and adapt to evolving system behaviors. As applications become increasingly complex and distributed, the challenge of accurately interpreting telemetry data will only intensify. Integrating feedback loops from engineers into the AI models will be crucial for refining root cause assessments and ensuring that STAR remains a valuable asset. The question now is whether we will see more organizations adopt a similar build-vs-buy approach to observability, or if the market will ultimately coalesce around a few dominant AI-powered platforms. The upcoming KDD conference in Jeju, as mentioned in [Anyone heading to Jeju for KDD? Let’s meet up! 🙋[D]]( /post/anyone-heading-to-jeju-for-kdd-let-s-meet-up-d-cmrwr0pd206r3djxxdw0a7k38), will likely offer further insights into these trends.

Expedia Group has introduced STAR, an internal AI-assisted observability platform that helps engineers investigate production incidents using service telemetry and LLMs. Built with FastAPI, Datadog, Celery, Redis, and Langfuse, STAR follows structured workflows to analyze telemetry, generate root cause assessments, and support incident response while keeping engineers in the loop.
By Leela KumiliRead on the original site
Open the publisher's page for the full experience