incident response

How AI Helps Instacart Engineers Resolve Incidents Faster in Slack

When a production incident lands in Slack, every minute of investigation matters.

3 min readInfoQ
How AI Helps Instacart Engineers Resolve Incidents Faster in Slack

Instacart's Blueberry is a welcome sign of maturity in how we apply AI to operational reliability. Rather than promising to replace engineers, this system does something more practical: it helps them investigate production incidents faster by organizing context and generating grounded hypotheses. That distinction matters. Blueberry operates in Slack, pulls from operational data and historical incident knowledge, and uses parallel subagents to surface root cause possibilities without removing the engineer from the decision loop. It is a tool for augmentation, not automation. This approach stands in sharp contrast to recent incidents where AI agents operated without sufficient boundaries, as seen in the case where AI Agents Shared User Images, Highlighting Data Security Concerns. Blueberry keeps the human in control, which is the difference between a helpful assistant and a liability.

For on-call engineers, the practical value here is time. Incident response is often a frantic scramble through logs, dashboards, and memory. Blueberry compresses that investigation phase by letting an AI agent do the parallel work of checking known failure patterns and correlating signals. The system uses MCP integrations to connect with existing infrastructure, meaning it does not require a wholesale rewrite of how your team operates. It fits into the workflow you already have. If you have ever spent an hour chasing a red herring during an outage, you can see the appeal of a tool that says, "Here are three plausible causes, and here is why each one fits the evidence." The key constraint is that the engineer still validates and decides. That is the right balance. Too many AI tools promise to fix everything and end up creating more noise. Blueberry appears to aim for clarity, not certainty.

There is a broader implication here for how teams think about incident data. Blueberry's effectiveness depends on a well-maintained history of past incidents and their resolutions. If your team does not already document postmortems in a structured, accessible way, this tool will not magically fix that. The quality of the hypotheses depends on the quality of the history. This is a reminder that AI-assisted operations are not a shortcut around good engineering hygiene; they are a reason to invest in it. Meanwhile, the security and data governance questions that arise with any AI agent connected to production systems cannot be ignored. The Protecting Your Data: Kiteworks Advises Temporary Server Shutdown story underscores that even established platforms face credible threats when handling sensitive data. Instacart will need to be transparent about how Blueberry isolates incident data and prevents leakage across customers or environments.

The specific detail to watch is how Blueberry handles ambiguous incidents where the root cause does not match any historical pattern. The system is grounded in precedent, which is powerful for common failure modes but could leave engineers blind to novel issues. Instacart's documentation of how Blueberry communicates uncertainty, or explicitly says "I do not know", will determine whether engineers trust it or ignore it. That is the difference between a tool that saves time and one that wastes it.

From InfoQ

Instacart introduced Blueberry, an AI-assisted incident response system that helps on-call engineers investigate production issues faster. It combines AI agents, operational data, and historical incident knowledge to generate grounded root cause hypotheses in Slack. It uses parallel subagents, MCP integrations, and incident history to reduce investigation time while keeping engineers in control.

Read the original at InfoQ