Self-RAG-AI-Agent: Self-RAG AI Agent: The Retrieval Loop That Checks Its Own Work
A LangGraph-powered document Q&A system that rewrites weak queries, verifies grounding, and only answers when the context actually supports it.
- This repository turns Self-RAG into a cautious control loop that refuses to answer until retrieval, grounding, and usefulness all pass.
- Its value is not just better prompting, but explicit stateful branching in LangGraph with rewrite attempts when support is weak.
- The project is strongest as a practical blueprint for enterprise document Q&A, where a wrong answer is more expensive than a slow one.
- The Streamlit frontend makes the agent legible by exposing execution details instead of hiding the verification steps.
The chatbot that won’t answer too early
Most RAG apps retrieve, generate, and hope for the best. This one does something closer to editorial review. It checks whether the retrieved context is relevant, whether the answer is actually supported, and whether the final response is useful enough to send.
That difference matters in internal document Q&A. Policies, pricing sheets, and company handbooks punish confident mistakes. This repo treats that as the core design constraint, not a cleanup task.
Implementing Self-RAG with LangGraph, LangChain, and Ollama.
Why Self-RAG is different from standard RAG
Standard RAG is linear. A question becomes a retrieval query, the vector store returns matches, and the model answers from whatever context it got. If the retrieval is weak, the answer can still sound polished.
Self-RAG adds judgment. In this repo, the model can critique the retrieved docs, critique its own generated answer, and decide to rewrite the query instead of pretending the first pass was good enough.
| Pattern | Retrieval | Validation | Rewrite loop | Best fit |
|---|---|---|---|---|
| Standard RAG | One pass | Usually none | No | General chat and lightweight assistants |
| Self-RAG AI Agent | Conditional | Support and usefulness checks | Yes | Enterprise document Q&A |
| Canonical Self-RAG paper | Reflective retrieval and critique | Model-centric reflection | Yes | Research and advanced implementations |
| Corrective RAG style systems | Retrieval quality focused | Pre-generation correction | Sometimes | Noisy corpora and search fallback workflows |
The agent state machine under the hood
The heart of the repo is the state object in backend.py. It tracks the question, retrieval query, documents, support verdicts, usefulness verdicts, and rewrite attempts. That state is what lets the graph branch instead of marching forward blindly.
class State(TypedDict):
question: str
retrieval_query: str
docs: List[Document]
issup: Literal["fully_supported", "partially_supported", "no_support"]
isuse: Literal["useful", "not_useful"]
rewrite_tries: int
# ... additional tracking fields
The interesting part is not that LangGraph runs steps. It is that each step can emit a verdict. Those verdicts change the path. If support fails, the graph can rewrite the query and retrieve again instead of forcing a final answer.
The three gates
The repo’s self-checks separate three jobs that often get conflated. Relevance asks whether the retrieved context is worth keeping. Support asks whether the answer is grounded in that context. Usefulness asks whether the answer actually solves the user’s task.
| Gate | What it checks | What happens if it fails |
|---|---|---|
| Relevance | Whether retrieved documents match the question | Discard weak context and keep filtering |
| Support | Whether the answer is grounded in the selected context | Rewrite the query and retrieve again |
| Usefulness | Whether the response is actually helpful to the user | Refine or retry before final output |
From notebooks to a working app
The notebook sequence reads like a deliberate build-up: start with retrieval, add context state, add support verification, then graduate to streaming and hosted models. That progression is useful because it shows the system as something engineered, not just declared.
This matters for readers because it lowers the intimidation factor. The repo is not asking you to swallow a giant agent framework in one leap. It shows the working parts one layer at a time.
The frontend makes the agent legible
The Streamlit app does more than stream tokens. Its execution details panel exposes the internal markers that make the system trustworthy: rewrite tries, support status, and other state transitions. That is good product design because it explains the answer, not just the answer text.
class StreamHandler(BaseCallbackHandler):
def on_llm_new_token(self, token: str, **kwargs) -> None:
# Stream tokens into the UI as they are generated
...
with st.expander("Execution Details"):
st.write(f"Rewrite Tries: {state['rewrite_tries']}")
st.write(f"IsSUP: {state['issup']}")
st.write(f"IsUSE: {state['isuse']}")
Where this sits in the RAG landscape
Compared with raw Self-RAG research code, this repo is easier to approach because it uses LangChain and LangGraph instead of forcing developers to start from a training-centric implementation. Compared with generic RAG templates, it adds explicit self-critique and retry logic. Compared with corrective RAG patterns, it leans harder into reflection as part of the answer path.
| Lens | Standard RAG | This repo |
|---|---|---|
| Control flow | Linear | Conditional and looping |
| Confidence handling | Implicit | Explicit support and usefulness checks |
| Query quality | Assumed | Rewritten when weak |
| Operational fit | Broad but brittle | Better suited to document-heavy enterprise Q&A |
Why this matters for enterprise document Q&A
The business case is simple. If the system answers a pricing question badly, a policy question badly, or an internal handbook question badly, the loss is not cosmetic. This design reduces that risk by favoring caution, traceability, and retry logic over speed alone.
That makes the project more than a demo. It is a practical pattern for teams that want LLM systems to behave less like improv and more like a disciplined analyst.