Retrieval-Augmented Generation (RAG)

Build and benchmark production RAG with fully open models

Building a RAG system that actually works in production takes more than stitching together a vector database.

3 min readMachine Learning

Most teams building retrieval-augmented generation today are flying blind. They bolt a vector database onto an LLM, run a few hand-picked examples, and call it production-ready. That approach breaks down the moment real-world queries arrive with typos, synonyms, or ambiguous phrasing. This workshop on August 29, led by Ben Auffarth of Chelsea AI Ventures, takes a different route: build a RAG pipeline entirely from open models, no API calls, and benchmark it end to end. That distinction matters. It is one thing to demo a chat interface over your own documents; it is another to measure whether the system actually retrieves the right chunks, reranks them effectively, and stays within budget.

The agenda reads like a checklist of what we wish every RAG team would do before shipping. Hybrid retrieval, combining vector search with keyword matching, acknowledges that dense embeddings alone miss exact-match terms and rare identifiers. Reranking then catches the relevant chunks that initial recall left on the table. Evaluation with RAGAS means quality changes are measured, not assumed, which is the only honest way to iterate. And guardrails designed in from the start, rather than bolted on after an incident, reflect a mature understanding of deployment risk. This is not flashy, but it is the difference between a demo and a system. We have been tracking how practical AI evaluation is evolving, from Jev vs LLMs: Evaluating AI for Practical Decision-Making to the broader deployment questions raised at Explore the Future of AI Deployment: Key Topics at QCon AI New York. This workshop sits squarely in that conversation, but with a hands-on twist: participants build and benchmark, not just listen.

What stands out is the cost and performance benchmarking for open-model deployments. Most content in this space tells you what to build, not what it costs to run or how fast it responds. By publishing actual numbers, this session gives attendees a baseline they can take back to their own projects. That is rare and valuable. If a reader asked us whether this is worth their time, the answer is yes, but with a caveat: go in ready to measure your own assumptions. The workshop will not hand you a magic pipeline, it will give you a methodology. In a field where hype outpaces evidence, that is the most practical thing on offer.

The open-model constraint is also worth noting. No API calls means no vendor lock-in, no per-token surprise bills, and full control over your data. That aligns with the broader push toward Scale Sandboxes Instantly: A New Approach to Concurrent AI Workloads, where teams are looking for infrastructure that does not force them into a corner. The takeaway here is specific: if you are serious about production RAG, stop treating retrieval as an afterthought. Start with hybrid search, add reranking, measure with RAGAS, and budget for guardrails. The details will be on the table August 29. The question is whether you will show up with your own benchmarks to challenge.

From Machine Learning

There’s a hands-on workshop on August 29 that builds and benchmarks this properly, end to end, using entirely open models, no API calls involved. Led by Ben Auffarth, AI Consultant and Founder of Chelsea AI Ventures.

• Hybrid retrieval (vector + keyword, not vector alone) • Reranking to catch relevant chunks that vector search alone misses • Evaluation with RAGAS, so quality changes are measured, not assumed • Guardrails built in from the design stage • Actual cost and performance benchmarking for open-model deployments

Read the original at Machine Learning