Self-RAG-AI-Agent: Self-RAG AI Agent: The Retrieval Loop That Checks Its Own Work

A LangGraph-powered document Q&A system that rewrites weak queries, verifies grounding, and only answers when the context actually supports it.

8 min read • View on GitHub • More from ASHishYADAav2003

A wide black-ink editorial illustration of a machine processing a question and a stack of documents through three inspection chambers. One chamber sends weak output back into a rewrite chute instead of releasing it. It explains that this agent treats answering as a gated process, not a single pass.
The core idea is not retrieval. It is refusal: the system checks, rejects, and retries before it answers.
Key Takeaways

The chatbot that won’t answer too early

Most RAG apps retrieve, generate, and hope for the best. This one does something closer to editorial review. It checks whether the retrieved context is relevant, whether the answer is actually supported, and whether the final response is useful enough to send.

That difference matters in internal document Q&A. Policies, pricing sheets, and company handbooks punish confident mistakes. This repo treats that as the core design constraint, not a cleanup task.

Implementing Self-RAG with LangGraph, LangChain, and Ollama.

Ashish Yadav, Project Author · Self-RAG-AI-Agent GitHub README

Why Self-RAG is different from standard RAG

Standard RAG is linear. A question becomes a retrieval query, the vector store returns matches, and the model answers from whatever context it got. If the retrieval is weak, the answer can still sound polished.

Self-RAG adds judgment. In this repo, the model can critique the retrieved docs, critique its own generated answer, and decide to rewrite the query instead of pretending the first pass was good enough.

PatternRetrievalValidationRewrite loopBest fit
Standard RAGOne passUsually noneNoGeneral chat and lightweight assistants
Self-RAG AI AgentConditionalSupport and usefulness checksYesEnterprise document Q&A
Canonical Self-RAG paperReflective retrieval and critiqueModel-centric reflectionYesResearch and advanced implementations
Corrective RAG style systemsRetrieval quality focusedPre-generation correctionSometimesNoisy corpora and search fallback workflows

The loop is the story: the agent can route back to rewrite the question when support is weak, instead of shipping a shaky answer.


The agent state machine under the hood

The heart of the repo is the state object in backend.py. It tracks the question, retrieval query, documents, support verdicts, usefulness verdicts, and rewrite attempts. That state is what lets the graph branch instead of marching forward blindly.

class State(TypedDict):
    question: str
    retrieval_query: str
    docs: List[Document]
    issup: Literal["fully_supported", "partially_supported", "no_support"]
    isuse: Literal["useful", "not_useful"]
    rewrite_tries: int
    # ... additional tracking fields

The interesting part is not that LangGraph runs steps. It is that each step can emit a verdict. Those verdicts change the path. If support fails, the graph can rewrite the query and retrieve again instead of forcing a final answer.

The three gates

The repo’s self-checks separate three jobs that often get conflated. Relevance asks whether the retrieved context is worth keeping. Support asks whether the answer is grounded in that context. Usefulness asks whether the answer actually solves the user’s task.

GateWhat it checksWhat happens if it fails
RelevanceWhether retrieved documents match the questionDiscard weak context and keep filtering
SupportWhether the answer is grounded in the selected contextRewrite the query and retrieve again
UsefulnessWhether the response is actually helpful to the userRefine or retry before final output
A close-up black-ink illustration of a sentence splitting into two paths. One path goes into a messy shelf of document fragments and returns scattered context cards. The other path gets rewritten into a cleaner query and returns a tighter stack of matched cards with a support stamp above them. It explains why query rewriting improves retrieval quality before generation.
Weak natural-language questions often miss the structure of the indexed documents. Rewrite the query first, and retrieval gets much cleaner.

From notebooks to a working app

The notebook sequence reads like a deliberate build-up: start with retrieval, add context state, add support verification, then graduate to streaming and hosted models. That progression is useful because it shows the system as something engineered, not just declared.

This matters for readers because it lowers the intimidation factor. The repo is not asking you to swallow a giant agent framework in one leap. It shows the working parts one layer at a time.

The frontend makes the agent legible

The Streamlit app does more than stream tokens. Its execution details panel exposes the internal markers that make the system trustworthy: rewrite tries, support status, and other state transitions. That is good product design because it explains the answer, not just the answer text.

class StreamHandler(BaseCallbackHandler):
    def on_llm_new_token(self, token: str, **kwargs) -> None:
        # Stream tokens into the UI as they are generated
        ...

with st.expander("Execution Details"):
    st.write(f"Rewrite Tries: {state['rewrite_tries']}")
    st.write(f"IsSUP: {state['issup']}")
    st.write(f"IsUSE: {state['isuse']}")

Where this sits in the RAG landscape

Compared with raw Self-RAG research code, this repo is easier to approach because it uses LangChain and LangGraph instead of forcing developers to start from a training-centric implementation. Compared with generic RAG templates, it adds explicit self-critique and retry logic. Compared with corrective RAG patterns, it leans harder into reflection as part of the answer path.

LensStandard RAGThis repo
Control flowLinearConditional and looping
Confidence handlingImplicitExplicit support and usefulness checks
Query qualityAssumedRewritten when weak
Operational fitBroad but brittleBetter suited to document-heavy enterprise Q&A

Why this matters for enterprise document Q&A

The business case is simple. If the system answers a pricing question badly, a policy question badly, or an internal handbook question badly, the loss is not cosmetic. This design reduces that risk by favoring caution, traceability, and retry logic over speed alone.

That makes the project more than a demo. It is a practical pattern for teams that want LLM systems to behave less like improv and more like a disciplined analyst.