Adaptive-RAG: The Repo That Decides When Not to Retrieve

A LangGraph-powered AI system that routes each question to local search, general model knowledge, or web search, then checks its own answers before they reach the user.

8 min read • View on GitHub • More from umang-bhavsar

A wide editorial scene of a routing switchyard with one central lever splitting a question into three paths. One path leads to stacked local documents, one to a model chamber, and one to a web search tower. A second inspector figure stands at the end with a magnifying glass, showing that answers are checked before they ship. This image explains that the system treats retrieval as a routing decision, not a default step.
Adaptive-RAG starts by choosing a path, then verifies the answer before it returns anything to the user.
Key Takeaways

Most RAG systems assume retrieval is always the right first move. This repo makes a sharper bet: first classify the question, then decide whether to retrieve from local docs, ask the model directly, or call the web. That changes the project from a chatbot wrapper into a control system for answer paths.

A novel adaptive QA framework that can dynamically select the most suitable strategy for (retrieval-augmented) LLMs from the simplest to the most sophisticated ones based on the query complexity.

Soyeong Jeong, Lead Author of the Adaptive-RAG paper · Adaptive-RAG: Official Code Repository

Why this RAG system starts by asking whether it should retrieve at all

The first surprise in umang-bhavsar/adaptive-rag is simple: retrieval is optional. A query classifier sends each request down one of three lanes, and only one of them hits local documents. The others go straight to general model knowledge or out to web search.

That design matters because it removes the default tax most RAG apps pay on every question. Easy questions do not need a vector lookup. Fresh questions do not need a stale local corpus. Mixed workloads need a policy, not a single retrieval habit.

A close-up mechanical loop on a drafting table. A query card enters a classifier slot, then a handoff arm drops it into one of three trays. The answer is placed under a red inspection lamp, and when it fails the check, a second arm rewrites the card and sends it back through the loop. This image explains the system's grade-and-rewrite behavior.
The most interesting part is not retrieval itself. It is the rewrite loop that repairs weak answers and tries again.

The control loop: classify, retrieve, grade, rewrite

The graph is not a straight pipeline. It is a policy engine with branches, a verifier, and a retry path.

# src/rag/graph_builder.py, simplified
route = query_classifier(user_query)

if route == "index":
    context = retriever(user_query)
elif route == "general":
    context = general_llm(user_query)
else:
    context = web_search(user_query)

answer = generate(context, user_query)
faithful = verify_answer(answer, context)

if not faithful and rewrite_attempts < MAX_REWRITE_ATTEMPTS:
    user_query = rewrite(user_query)
    return run_graph(user_query)

return answer

That loop is the repo's real center of gravity. It does not just retrieve. It grades the output, and if the answer is not supported, it rewrites the query and tries again. That makes the system feel less like a prompt chain and more like a small fault-tolerant machine.

The code path is shaped by src/rag/graph_builder.py, which behaves like a state machine rather than a linear script. The key move is the handoff from classification to routing. After that, the verifier acts like a referee, not a decorative post-processing step.

Why the repo feels production-minded

The architecture is not just clever. It is practical. FastAPI handles the backend, Streamlit handles the interface, MongoDB keeps chat history, and the vector store stays isolated inside retrieval setup code. That separation matters when a demo starts behaving like a real product.

Persistent memory is the quiet tell. MongoDBChatMessageHistory means the system can remember a session across restarts instead of pretending context lives only in RAM. The prompt layer is also decoupled, so behavior can change without rewriting the orchestration logic.

# src/config/prompts.yaml
classify_prompt: |
  Decide whether the question should go to local docs, general knowledge, or web search.

verify_prompt: |
  Check whether the answer is supported by the retrieved context.
  Return a judgment, not a rewrite.

rewrite_prompt: |
  Rewrite the query to improve retrieval quality without changing intent.

The prompt layer is separate on purpose

That YAML file is not a convenience. It is an operating principle. In agent systems, prompts tend to drift faster than code, so pulling them out of Python keeps the system editable without making the graph brittle.

Adaptive-RAG is a really nice approach that brings together some interesting ideas... of flow engineering or adaptive RAG which is the idea of doing tests in your rag pipeline or in your rag inference flow to do things like check relevance documents.

Lance Martin, Developer at LangChain · LangGraph: Adaptive RAG

What makes this different from naive RAG, Self-RAG, and GraphRAG

SystemRetrieval policySelf-checkingLatency profileBest use caseMain cost
Naive RAGAlways retrieveUsually nonePredictable, but wasteful on easy questionsSimple document QAUnnecessary retrieval on every query
Self-RAGRetrieves and reflects during generationYesHigher, because of more reflection cyclesComplex reasoning and critique-heavy tasksMore compute and more moving parts
CRAGRetrieves, then grades retrieved contentYes, before generationModerate to highNoisy retrieval environmentsExtra grading step before answering
GraphRAGPreprocesses into graph structure, then retrieves relationallyIndirectlyFront-loaded preprocessing costEntity-rich, relational corporaHeavy indexing and setup cost
Adaptive-RAGChooses whether to retrieve at all, then verifies the answerYes, through a closed loopLower on easy questions, higher only when neededMixed-complexity workloadsRouting logic plus retry orchestration

The key difference is not sophistication for its own sake. It is restraint. Adaptive-RAG is trying to spend the least amount of compute that still gets a trustworthy answer. That is a different goal from maximizing retrieval everywhere.

The trade-off: smarter routing, more moving parts

The downside is obvious once you see the upside. A routing system is harder to debug than a fixed pipeline. Each branch introduces its own failure mode, and the verifier can hide problems if you do not inspect the intermediate state carefully.

But that complexity is the point. Production AI systems are not clean lines from prompt to answer. They are choices, fallbacks, checks, and repairs. This repo is honest about that reality.

The bigger lesson

The broader lesson is not just that LangGraph is useful. It is that retrieval should be treated as a decision, not a reflex. Good systems should know when to spend tokens, when to stay local, and when to say the answer is not yet good enough.

That is why adaptive-rag stands out. It is not a retrieval system with extras. It is a routing system for uncertainty.