enterprise data management

Hybrid data queries demand agents that can adapt, not just stronger models.

Databricks' latest research reveals that traditional single-turn retrieval-augmented generation (RAG) systems struggle with hybrid queries, particularly when combining structured and unstructured data.

3 min readVentureBeat
Hybrid data queries demand agents that can adapt, not just stronger models.

The takeaway here is straightforward: if your data questions live on one side of the structured-unstructured divide, single-turn RAG will keep failing you, and no amount of model upgrading will fix that. Databricks's research makes the case that the bottleneck is architectural, not a matter of raw intelligence. A stronger foundation model still lost to the multi-step agent by 21% on the academic domain and 38% on the biomedical domain. That gap is not a rounding error. It is evidence that the way the system is built, not the model it runs on, determines whether hybrid queries get answered.

For data teams, this changes the calculus on where to invest engineering effort. The instinct to keep tuning a custom RAG pipeline, flattening SQL tables into text chunks, normalizing JSON into embeddings, rewriting retrieval logic for each new source, is a treadmill. Databricks's Supervisor Agent approach treats the problem differently. It lets the agent issue parallel calls to SQL and vector search, detect when results do not overlap, and issue a follow-up query like a SQL JOIN without requiring the data to be pre-merged. The practical effect is that adding a new data source becomes a configuration task, writing a plain-language description of what it contains, rather than a conversion project. That is the difference between maintaining a pipeline and describing an intent.

The research also draws a useful line in the sand about what this is not. This is not a hybrid retrieval technique where you blend embeddings with table results and hope for the best. It is an agent with access to multiple tools, and that distinction matters because it scales. Hybrid retrieval forces you to anticipate every kind of join a user might ask for. An agent can reason about the join on the fly, and it can correct itself when the first pass comes up empty. That self-correction is not a nice-to-have; it is the difference between a system that answers a question about declining sales by finding reviews that mention the product, and one that verifies the product actually had declining sales by checking the table first.

The limits are real, and they are worth respecting. Five to ten data sources is a manageable range. Connect too many at once and the agent slows down and starts routing to contradictory sources. And no architecture can fix data that is factually wrong. But the direction is clear. If your roadmap involves questions that span sales figures and customer reviews, or citation counts and academic papers, you are not going to close that gap by waiting for a better model. You need an agent that can bring the query to the data, not the other way around. Build for that, and the 20% gains on benchmarks start to look like a floor, not a ceiling.

From VentureBeat

Data teams building AI agents keep running into the same failure mode. Questions that require joining structured data with unstructured content, sales figures alongside customer reviews or citation counts alongside academic papers, break single-turn RAG systems.

Read the original at VentureBeat