Search Engine

Explore how hybrid search transforms discovery on Papers with Code

Building a better search engine for research papers rarely comes down to a single clever trick.

3 min readMachine Learning
Explore how hybrid search transforms discovery on Papers with Code
How we built a SOTA search engine using PostgreSQL, pgvector, and Qwen3 embeddings [P]

The most interesting thing about the search engine powering Papers with Code is not the novelty of the stack. It is the quiet confidence in using PostgreSQL, pgvector, and Qwen3 embeddings to blend keyword and semantic search. The approach is refreshingly straightforward, and that is exactly why it works. It does not claim to have invented a new paradigm. It simply demonstrates that a hybrid approach, combining lexical precision with semantic understanding, outperforms either method in isolation. This is the kind of practical engineering that moves the field forward, not by chasing the latest framework, but by optimizing the tools we already have.

For our readers who are wrestling with their own search problems, this is a masterclass in pragmatic architecture. The decision to use a live embedding model served through Hugging Face Inference Endpoints, alongside batch generation via Jobs, highlights a mature understanding of cost and latency trade-offs. It is easy to get lost in the weeds of model selection or vector database hype. But this breakdown brings us back to the fundamentals: a reliable relational database, a well-supported extension, and a clear-eyed view of what you are trying to achieve. This is a direct counterpoint to the more abstract discussions we often see about Unlock Advanced RAG: 6 Architectures for Semantic Search & LLMs. That piece explores the broader design space, but here we see a concrete, working example. The value is in the specifics, from the choice of a 0.6B parameter model to the use of a standard PostgreSQL instance. It proves that you do not need a bespoke infrastructure to deliver meaningful improvements.

What is also notable is the implicit commentary on the current state of AI infrastructure. The fact that the same system powers the "related papers" recommendations shows a deliberate effort to get more value from a single investment. This is not an experiment; it is production infrastructure. It also speaks to a broader trend we have touched on before, where the real challenges are not about the AI model itself, but about the surrounding plumbing. This is not a story about a model "escaping" its constraints, as we see in Beyond the Hype: Why AI "Escapes" Are Really Firewall Shortcomings. Instead, it is a story about building reliable, constrained systems that do exactly what they are asked to do.

The open question this raises for the community is about the ceiling of this approach. The team is candid about the hybrid strategy producing better results than either method alone, but how much headroom is left? As embedding models become more powerful and context-aware, will the semantic component start to outweigh the keyword side? And for those of us working with technical content, is there a point where the noise from semantic matching begins to hurt precision? We would argue that the next breakthrough will not come from a new model architecture, but from smarter ways to combine and rerank these signals. The specific takeaway here is simple: start with what you have, add a good embedding model, and measure the impact. The infrastructure is ready. The question is whether you are ready to build on it.

From Machine Learning

I wrote a technical breakdown of how search works on Papers with Code.

The system combines keyword and semantic search, which produced better results than either approach alone. The stack includes:

Read the original at Machine Learning