generative AI for data analysis

Give your AI agents the data context they need to find the right answer

SQL query logs are crucial for AI agents to avoid misinterpreting data joins, as demonstrated by Miro's experience with over 10,000 tables in Snowflake, where inaccuracies arose more than 65% of the time.

3 min readVentureBeat
Give your AI agents the data context they need to find the right answer

In a landscape where data-driven decision-making is paramount, Miro’s experience with AI agents in its Snowflake environment offers a critical lesson about the importance of context in data management. Initially, Miro's data team discovered that when these agents were pointed directly at over 10,000 tables, they yielded the wrong answers more than 65% of the time. This wasn't due to the AI model's inadequacy; rather, it stemmed from a lack of contextual understanding. Without a semantic layer to guide the agents, they struggled to identify which data assets corresponded to specific business questions. This scenario underscores a broader challenge faced by organizations attempting to leverage AI for analytics—a topic that resonates well with trends in the industry, such as those discussed in How DeepSeek’s radical architecture is shattering Silicon Valley's token moat and Has the hunt for AI compute uncovered the next Cerebras?.

The recent introduction of DataHub's Context Intelligence layer, which leverages existing SQL query history to build a semantic index, marks a significant advancement in addressing these challenges. By transforming raw query logs into a structured knowledge base, DataHub empowers AI agents to access previously validated queries, thereby reducing the risk of erroneous data interpretations. This capability is particularly profound as it shifts the paradigm from merely cataloging data to creating a dynamic, context-rich environment that enhances the accuracy of AI responses. As organizations increasingly rely on automated agents for data queries, the ability to mine and utilize historical context will be crucial in ensuring reliable outcomes.

The essence of DataHub's innovation lies not just in its technical sophistication but also in its human-centered approach. By allowing domain experts to validate AI-generated context, DataHub bridges the gap between raw data and actionable insights. This validation process ensures that discrepancies in metric calculations across different teams are identified and addressed, promoting a unified understanding of data within the organization. The implications for enterprises are substantial; as more organizations adopt AI-driven data solutions, the ability to contextualize data effectively will distinguish leaders in the field from those who struggle with data accuracy and reliability. This transition resonates with the evolving role of designers in tech, as noted in articles like Are designers the new SWEs? Figma Make's new two-way GitHub integration turns designs into live, production code — with built-in governance.

Looking ahead, the emergence of context layers such as DataHub’s could redefine how organizations interact with their data assets. As the competition for context management intensifies, it raises critical questions about who will control the narrative around data-driven decisions. Enterprises must remain vigilant, ensuring their context management solutions are not only comprehensive but also compatible with their existing infrastructures. The shift from passive data catalogs to continuously updated semantic intelligence could very well become the next frontier in data management, significantly influencing how organizations leverage AI to enhance their operational efficiency and decision-making processes. The strategic importance of context in AI applications is now clearer than ever, setting the stage for an exciting evolution in how businesses harness the power of their data.

From VentureBeat

When Miro’s data team pointed AI agents directly at its Snowflake environment, the agents got the wrong answer more than 65% of the time. The problem wasn’t the model — it was context. With more than 10,000 tables and no semantic layer to guide routing, the agents had no way to know which data assets matched which business questions.

Read the original at VentureBeat