production-agentic-rag-course: The RAG Course That Treats Search Like a System, Not a Demo
A production-first curriculum for building an arXiv research assistant with BM25, hybrid retrieval, LangGraph feedback loops, Airflow ingestion, and local model infrastructure.
- This course treats retrieval as the product and the prompt as a thin layer on top.
- Its real differentiator is a failure-aware loop that rewrites queries when search quality is weak.
- The stack is deliberately full, from Airflow and OpenSearch to LangGraph and Langfuse, because production RAG is mostly systems work.
- The curriculum reads like an implementation roadmap for builders who need a real research assistant, not a notebook demo.
Most RAG tutorials start with embeddings and end with a confident answer. This repo starts somewhere more honest: search. For academic content, exact terms, paper IDs, and domain jargon often matter more than a clever prompt, and this course treats that as the default, not the exception.
Why this repo starts with search, not embeddings
The opening move is blunt. Build BM25 first. Then add vectors. Then layer on agentic orchestration. That order is the point, because technical research systems fail when they act like semantic similarity can replace precise retrieval.
| Naive RAG | This repo | Why it matters |
|---|---|---|
| Vector DB plus prompt | BM25 plus vectors plus reranking | Exact matches survive, so technical queries stay grounded. |
| One retrieval pass | Search, grade, rewrite, then answer | The system can recover when the first query misses. |
| Thin notebook prototype | Dockerized services with Airflow, OpenSearch, Postgres, Redis | The course teaches production plumbing, not just model calls. |
| Answer-first thinking | Search-first thinking | The retrieval layer becomes the product, not a detail. |
| Best for broad prose | Best for technical papers and messy queries | Academic language rewards keyword precision and hybrid ranking. |
That difference shows up in the course’s structure. It does not pretend that vector search is enough. It teaches the boring parts, because the boring parts are what make the system useful.
The self-correcting loop that changes the game
The central idea lives in agentic_rag.py. A LangGraph state machine retrieves documents, grades them, and only generates an answer if the evidence is good enough. If the retrieval is weak, the query gets rewritten and sent back through the search path.
We build RAG systems the way successful companies do - solid search foundations enhanced with AI, not AI-first approaches that ignore search fundamentals.
That quote is the project in one sentence. The agent is not a replacement for search. It is a correction layer on top of search, which is a much more realistic way to think about production research assistants.
What the stack is doing behind the curtain
The infrastructure is not decoration. FastAPI exposes the system, OpenSearch handles both BM25 and vector retrieval, PostgreSQL keeps metadata honest, Redis absorbs speed work, Airflow manages ingestion, Langfuse watches the traces, Docling parses the PDFs, and Ollama keeps the model layer local.
| Layer | What it does | Why it matters |
|---|---|---|
| FastAPI | Serves search and agent endpoints | Makes the system testable and modular. |
| OpenSearch | Runs BM25 and vector retrieval | Keeps retrieval native to the search engine. |
| PostgreSQL | Stores metadata | Separates structured state from retrieved content. |
| Redis | Caches hot data | Protects latency when the stack grows. |
| Airflow | Schedules ingestion | Turns paper collection into a repeatable pipeline. |
| Langfuse | Tracks traces | Makes failures visible instead of mysterious. |
| Docling | Parses academic PDFs | Handles layouts that simpler parsers break on. |
| Ollama | Runs local models | Supports a privacy-conscious, self-contained setup. |
This is why the repo feels more like a production system than a course repo. The curriculum teaches the scaffolding that surrounds retrieval, because that scaffolding is what keeps the assistant reliable when the dataset, the query, or the model changes.
Why hybrid search beats just vectors here
Hybrid retrieval is the obvious next move once you accept that search is the product. BM25 catches exact terms. Vector search catches meaning. Reciprocal Rank Fusion blends both so the system does not collapse when the query is technical, shorthand-heavy, or messy.
| BM25 | Vector search | Hybrid retrieval |
|---|---|---|
| Best at exact terms and paper IDs | Best at semantic similarity | Best when both kinds of matches matter. |
| Weak on paraphrase | Weak on literal precision | Reduces blind spots in technical search. |
| Simple to reason about | Flexible but fuzzy | Balances interpretability and recall. |
| Good default for academic text | Good default for concept matching | Good default for production RAG over papers. |
The important thing is not that hybrid search is fashionable. It is that it is resilient. When the user asks a precise question about a paper title, a method name, or a formula, the system still has a way to land on the right document.
A curriculum that reads like an implementation roadmap
The seven-week structure is itself an argument. First come infrastructure and ingestion. Then search. Then RAG. Then monitoring. Then the agent. That progression mirrors how serious systems get built in the wild.
| Curriculum phase | Builds what | What it teaches |
|---|---|---|
| Weeks 1-2 | Docker, FastAPI, Postgres, OpenSearch, Airflow | The system has to exist before the model can matter. |
| Weeks 3-4 | BM25 and hybrid search | Retrieval quality is a design problem, not an afterthought. |
| Weeks 5-6 | Full RAG, streaming, monitoring | Answers need observability and latency discipline. |
| Week 7 | LangGraph agentic RAG and bot interface | The system learns to judge and reroute itself. |
That makes the repo unusually valuable as a portfolio project. It does not just prove that someone can call an API. It shows that they can assemble a production-shaped stack and understand why each layer exists.
Who this course is really for
This is for builders who are tired of toy demos. AI engineers, full-stack developers, and technically curious practitioners will get the most out of it, especially if they care about search quality, observability, and deployability.
| Beginner demo mindset | Production builder mindset | Who wins |
|---|---|---|
| “Can I get an answer?” | “Can the system find the right evidence?” | The production builder. |
| Single prompt chain | Stateful retrieval loop | The system that can recover from failure. |
| One-off notebook | Services, traces, ingestion, and ranking | The team that needs reliability. |
| Vector-first habit | Search-first discipline | Anyone working with technical documents. |
If you want a weekend prototype, this is heavier than you need. If you want a credible blueprint for a research assistant, it is right in the sweet spot.





