production-agentic-rag-course: The RAG Course That Treats Search Like a System, Not a Demo

A production-first curriculum for building an arXiv research assistant with BM25, hybrid retrieval, LangGraph feedback loops, Airflow ingestion, and local model infrastructure.

10 min read • View on GitHub • More from jamwithai

A wide editorial scene of a research desk where printed papers, search rails, and a control panel form one machine. The image explains that retrieval is the center of the system, and the agent only steps in when the search results look weak.
Search is not a supporting role here. It is the machine the rest of the course is built around.
Key Takeaways

Most RAG tutorials start with embeddings and end with a confident answer. This repo starts somewhere more honest: search. For academic content, exact terms, paper IDs, and domain jargon often matter more than a clever prompt, and this course treats that as the default, not the exception.

Why this repo starts with search, not embeddings

The opening move is blunt. Build BM25 first. Then add vectors. Then layer on agentic orchestration. That order is the point, because technical research systems fail when they act like semantic similarity can replace precise retrieval.

Naive RAGThis repoWhy it matters
Vector DB plus promptBM25 plus vectors plus rerankingExact matches survive, so technical queries stay grounded.
One retrieval passSearch, grade, rewrite, then answerThe system can recover when the first query misses.
Thin notebook prototypeDockerized services with Airflow, OpenSearch, Postgres, RedisThe course teaches production plumbing, not just model calls.
Answer-first thinkingSearch-first thinkingThe retrieval layer becomes the product, not a detail.
Best for broad proseBest for technical papers and messy queriesAcademic language rewards keyword precision and hybrid ranking.

That difference shows up in the course’s structure. It does not pretend that vector search is enough. It teaches the boring parts, because the boring parts are what make the system useful.

The self-correcting loop that changes the game

The central idea lives in agentic_rag.py. A LangGraph state machine retrieves documents, grades them, and only generates an answer if the evidence is good enough. If the retrieval is weak, the query gets rewritten and sent back through the search path.

The agent does not just answer. It checks whether retrieval is good enough, and if not, it rewrites the question and tries again.

A close-up of a grading chamber where a retrieved snippet is tested against a query card. One path continues to answer generation, while the other path loops back through a rewrite arm. The image explains that retrieval quality is checked before the system commits to an answer.
The most important decision is not generation. It is whether the retrieved evidence deserves to become the answer.

We build RAG systems the way successful companies do - solid search foundations enhanced with AI, not AI-first approaches that ignore search fundamentals.

Shirin Khosravi Jam, Sr. Data Scientist & AI Engineer / Co-creator · GitHub - jamwithai/production-agentic-rag-course

That quote is the project in one sentence. The agent is not a replacement for search. It is a correction layer on top of search, which is a much more realistic way to think about production research assistants.

What the stack is doing behind the curtain

The infrastructure is not decoration. FastAPI exposes the system, OpenSearch handles both BM25 and vector retrieval, PostgreSQL keeps metadata honest, Redis absorbs speed work, Airflow manages ingestion, Langfuse watches the traces, Docling parses the PDFs, and Ollama keeps the model layer local.

LayerWhat it doesWhy it matters
FastAPIServes search and agent endpointsMakes the system testable and modular.
OpenSearchRuns BM25 and vector retrievalKeeps retrieval native to the search engine.
PostgreSQLStores metadataSeparates structured state from retrieved content.
RedisCaches hot dataProtects latency when the stack grows.
AirflowSchedules ingestionTurns paper collection into a repeatable pipeline.
LangfuseTracks tracesMakes failures visible instead of mysterious.
DoclingParses academic PDFsHandles layouts that simpler parsers break on.
OllamaRuns local modelsSupports a privacy-conscious, self-contained setup.

This is why the repo feels more like a production system than a course repo. The curriculum teaches the scaffolding that surrounds retrieval, because that scaffolding is what keeps the assistant reliable when the dataset, the query, or the model changes.

Why hybrid search beats just vectors here

Hybrid retrieval is the obvious next move once you accept that search is the product. BM25 catches exact terms. Vector search catches meaning. Reciprocal Rank Fusion blends both so the system does not collapse when the query is technical, shorthand-heavy, or messy.

BM25Vector searchHybrid retrieval
Best at exact terms and paper IDsBest at semantic similarityBest when both kinds of matches matter.
Weak on paraphraseWeak on literal precisionReduces blind spots in technical search.
Simple to reason aboutFlexible but fuzzyBalances interpretability and recall.
Good default for academic textGood default for concept matchingGood default for production RAG over papers.

The important thing is not that hybrid search is fashionable. It is that it is resilient. When the user asks a precise question about a paper title, a method name, or a formula, the system still has a way to land on the right document.

A curriculum that reads like an implementation roadmap

The seven-week structure is itself an argument. First come infrastructure and ingestion. Then search. Then RAG. Then monitoring. Then the agent. That progression mirrors how serious systems get built in the wild.

Curriculum phaseBuilds whatWhat it teaches
Weeks 1-2Docker, FastAPI, Postgres, OpenSearch, AirflowThe system has to exist before the model can matter.
Weeks 3-4BM25 and hybrid searchRetrieval quality is a design problem, not an afterthought.
Weeks 5-6Full RAG, streaming, monitoringAnswers need observability and latency discipline.
Week 7LangGraph agentic RAG and bot interfaceThe system learns to judge and reroute itself.

That makes the repo unusually valuable as a portfolio project. It does not just prove that someone can call an API. It shows that they can assemble a production-shaped stack and understand why each layer exists.

Who this course is really for

This is for builders who are tired of toy demos. AI engineers, full-stack developers, and technically curious practitioners will get the most out of it, especially if they care about search quality, observability, and deployability.

Beginner demo mindsetProduction builder mindsetWho wins
“Can I get an answer?”“Can the system find the right evidence?”The production builder.
Single prompt chainStateful retrieval loopThe system that can recover from failure.
One-off notebookServices, traces, ingestion, and rankingThe team that needs reliability.
Vector-first habitSearch-first disciplineAnyone working with technical documents.

If you want a weekend prototype, this is heavier than you need. If you want a credible blueprint for a research assistant, it is right in the sweet spot.