Multi-Document-RAG-Chatbot-System-: Multi-Document-RAG-Chatbot-System: The Small RAG App That Rebuilds Its Own Memory

A Streamlit, LangChain, FAISS, and Groq stack that turns a folder of PDFs into a queryable second brain, then exposes the source chunks so you can verify every answer.

5 to 6 min read • View on GitHub • More from Ishita-Agrawal03

A folder of PDFs labeled docs/ feeds pages into a compact mechanical index cabinet, then selected chunks travel toward a chat window with a visible source snippet card beside it. The image explains that the app turns local documents into a searchable memory and keeps the evidence close at hand.
The repo’s main trick is not just retrieval. It rebuilds its memory from local PDFs and keeps the source trail visible.
Key Takeaways

The part most RAG demos skip: proving the answer came from the right PDF

Most RAG demos stop at “ask a question, get an answer.” This repo does something more useful. It reloads the PDFs in docs/, refreshes the vector index, runs a startup query, and exposes the retrieved chunks so the answer can be checked against the original material.

This chatbot system allows users to interact with multiple documents simultaneously, leveraging Retrieval-Augmented Generation (RAG) to provide accurate and context-aware responses.

Ishita Agrawal, Author/Maintainer · Multi-Document-RAG-Chatbot-System-

That makes the project feel less like a demo and more like a small trust layer. The user does not just get an answer. They get a path back to the source, which matters whenever a document system is being used for study notes, internal PDFs, or any other material where provenance is the whole point.

Why this stack is interesting: local embeddings, cloud reasoning

The architecture is a clean split. Embeddings stay local with sentence-transformers/all-MiniLM-L6-v2, while the question-answering model runs on Groq with llama-3.3-70b-versatile. That gives the repo a useful shape: cheap retrieval, stronger reasoning, and no need to host a giant model on the user’s machine.

The stack is simple, but the control flow matters. Local retrieval builds the evidence, then Groq turns that evidence into an answer.

That division also tells you what the repository is optimizing for. It is not chasing a fully local LLM stack. It is choosing the parts that benefit most from locality, then outsourcing the part where a large hosted model can add polish and fluency.


Inside app.py: one file that does ingestion, chunking, retrieval, and UI

The codebase is intentionally small. The main pipeline lives in app.py, which globs docs/*.pdf, parses each file with PyPDFLoader, injects source metadata, chunks text with RecursiveCharacterTextSplitter, and builds a FAISS index that gets saved locally.

That monolithic shape is not elegant in a framework sense, but it is excellent for learning. You can trace the path from PDF to chunk to embedding to retrieval without jumping across directories or abstractions. For a prototype, that is a feature.

for doc in loaded_docs:
    doc.metadata["source_file"] = pdf

text_splitter = RecursiveCharacterTextSplitter(
    chunk_size=1000,
    chunk_overlap=200
)
chunks = text_splitter.split_documents(loaded_docs)

vectorstore = FAISS.from_documents(chunks, embeddings)
vectorstore.save_local("vectorstore")

qa = RetrievalQA.from_chain_type(
    llm=llm,
    retriever=vectorstore.as_retriever(),
    return_source_documents=True,
)

The important part is not the syntax. It is the contract. Every chunk keeps a source reference, every answer can return source documents, and the saved index means the app can reload its memory instead of pretending the world starts fresh on every run.

The warm-up query is the weirdest and smartest line in the repo

A close-up mechanical relay switch sits between a vector store and an LLM, with a test token passing through a retrieval gate before a readiness lamp flips from unverified to ready. The image explains startup validation as a system check that confirms the retrieval and QA path works before a user arrives.
The startup query turns a hidden reliability check into a visible design choice.

This is the repo’s most memorable move. On startup, it runs a similarity search and a QA call with a test prompt, which means the app checks the full retrieval path before it accepts a real question. That is a tiny but valuable operational pattern.

It changes the meaning of “works.” The system is not merely configured. It has already verified that the vector store is live, the model can answer, and the source path still resolves. In a small prototype, that is the closest thing to a heartbeat monitor.

What this repo is, and what it is not

DimensionThis repoLlamaIndexLangChain
Primary roleA compact multi-document RAG templateA data framework focused on indexing and retrievalA general orchestration toolkit for LLM apps
ScopeSingle-purpose and easy to readBroader retrieval and indexing featuresBroadest surface area and integrations
UIStreamlit app includedUsually paired with another UI layerUsually paired with another UI layer
Best useLearning the end-to-end shape of a RAG appBuilding more advanced retrieval workflowsComposing chains, tools, and agents
TradeoffClarity over breadthSpecialization over simplicityFlexibility over minimalism

This matters because the repository is honest about its size. There is no attempt to become a framework. There is no packaging polish, no grand abstraction layer, and no claim that it can outdo the bigger projects. Its value is that it shows the minimum viable shape of a working multi-document RAG app.

That also means the rough edges are informative. Debug prints, a lightweight file layout, and a local-first vector store are all signs of a learning artifact. In this case, the roughness is part of the lesson.

Compared with the bigger RAG stacks

Against LlamaIndex, LangChain, and Haystack, the repo is not trying to win on capability. It wins on compactness. You can see the full system in one sitting, which is exactly what makes it useful for people who want to understand how these pieces fit before they commit to a larger framework.

ProjectStrengthWeaknessWhere this repo sits
Multi-Document-RAG-Chatbot-SystemEasy to audit, easy to copyLimited feature depthA small, opinionated implementation
LlamaIndexStrong retrieval and indexing primitivesMore framework surface to learnA better fit for advanced document workflows
LangChainHuge ecosystem and orchestration powerCan feel broad before it feels focusedThe repo already uses it as a building block
HaystackEnd-to-end RAG and NLP workflow supportHeavier than a toy templateA stronger production-oriented option

So the niche is clear. This repository is best understood as a readable blueprint, not a destination. If you want a production platform, the larger frameworks are the obvious next stop. If you want to understand the control flow, this is the quicker read.

The take-away: a second brain is only useful if it can show its receipts

The strongest lesson in this repo is not about embeddings or even about RAG. It is about visibility. The system rebuilds its memory from local PDFs, checks itself at startup, and returns the source chunk behind the answer. That combination turns a chatbot into something closer to an inspectable document system.

That is why the project works as an article subject. It makes the mechanics legible. The reader can trace the evidence, understand the tradeoff between local retrieval and cloud reasoning, and walk away with a concrete template for building a small but trustworthy knowledge base.