`retrieval-augmented-ai-system-for-intelligent-document-analytics`: The Local RAG Workbench That Treats PDFs, CSVs, and Audio as One Product

A small Streamlit app with a big idea: route unstructured and structured questions through the simplest possible path, keep the data local, and turn document analysis into a multi-modal workflow instead of a single chat box.

8 min read • View on GitHub • More from Sreshta21

A wide editorial scene shows a desk-sized analytics station split into two paths. One side handles a paper PDF stack and document question answering, while the other turns a CSV grid into a chart and then into audio output from the same interface. It explains that the project is not a single chatbot but a routed document analytics workbench.
One interface, three outputs. The repo’s core idea is to route questions to the right modality instead of forcing everything through one chat loop.
Key Takeaways

Most RAG apps promise to answer questions about documents. This one tries to do something narrower and more useful: treat documents, tables, charts, and spoken output as one workflow. That makes the repo feel less like a chatbot demo and more like a compact analytics station.

Why this is more than a document chatbot

The usual document assistant is one lane. Upload a PDF, ask a question, get text back. This repo widens the lane. It handles unstructured document retrieval and structured CSV analysis in the same Streamlit app, which matters because many real workflows move between the two constantly.

That hybrid design changes the product shape. Instead of asking users to decide whether they need a chat tool or a data tool, the interface decides for them. The result is a workbench with opinionated defaults, not a pile of features.

The real trick: a simple router with strong opinions

The most interesting code path is also the least glamorous. In `streamlit_app.py`, the app uses a keyword-based router to decide whether the user is asking for a plot or for document retrieval. Plot language goes to Matplotlib. Document language goes to FAISS-backed search and synthesis.

The app is not one monolithic agent. It is a deterministic router that chooses between retrieval, plotting, and reuse of cached state.

That is the repo’s best opinion. It avoids the temptation to hand every decision to an agent framework. Instead, it draws a hard line between two problem types and routes them with a rule you can explain in one sentence.

if any(word in user_prompt.lower() for word in ["plot", "graph", "chart"]):
    return make_chart(df)
else:
    return ask_pdf(user_prompt, db)

How the RAG pipeline is wired

The retrieval side is split cleanly between ingestion and answering. `vector_store.py` builds the index with a recursive text splitter, HuggingFace embeddings, and FAISS. `ask_pdf.py` performs similarity search, then passes the retrieved context to Ollama with a prompt that tells the model to admit uncertainty when the answer is not in the text.

That last instruction matters more than it sounds. A grounded system is not defined by vector search alone. It is defined by what happens when retrieval fails. Here, the answer is not forced. The model is told to stop and say it does not know.

A close-up pipeline scene shows a PDF page being cut into chunks, then filtered, then three relevant tiles landing in a tray beside a note that says the system should admit uncertainty. A second tray widens to hold more candidates, showing the fallback retrieval path. It explains how grounding and fallback search work together.
The pipeline favors small, explainable steps. It starts narrow, then expands retrieval only when the first pass looks weak.

The implementation is practical rather than fancy. Chunking is explicit, retrieval is bounded, and the fallback path widens the search only when needed. That is the opposite of magical. It is the point.

Why the system feels local, private, and fast

The stack leans heavily on local components: Ollama for inference, FAISS for vector storage, HuggingFace embeddings for representation, and local session state for reuse. Even the speed story is partly a locality story. Once the app has cached the database and dataframe, reruns do less work.

DimensionThis repoTypical cloud RAG app
InferenceLocal Ollama modelsRemote hosted API
EmbeddingsLocal HuggingFace modelManaged embedding service
StorageFAISS on diskCloud vector store
Privacy postureData stays on machineData leaves the machine
Operational burdenUser manages models and filesProvider manages more of the stack

The trade-off is clear. You get sovereignty and predictability, but you also inherit more setup and maintenance. This is not the frictionless path. It is the controlled path.

The repo is also an iteration log

The `CODES/` directory turns the project into a record of experimentation. `pdf_perfect.py` shows prompt tuning, numbered responses, and fallback retrieval logic that expands `k` when the first answer looks thin. That is useful because it captures how these systems actually get better: not with one breakthrough, but with a sequence of corrections.

if "not found" in response.lower() or len(response.split()) < 40:
    docs = db.similarity_search(query, k=12)
    prompt += "\nPlease answer in numbered steps."

That kind of code is easy to underestimate. It is not polished product code, but it is honest RAG engineering. It shows the moments where retrieval quality, answer length, and response structure all had to be nudged into place.

What it gets right, and where it stops short

The repo is strongest as a disciplined prototype. It is not trying to pretend it solved orchestration, security, or deployment at scale. It has hardcoded paths, duplicated app files, and local-only assumptions that would need work before broader use.

StrengthLimit
Clear routing between PDF and CSV tasksKeyword heuristics can be brittle
Local-first architectureMore setup for the user
Readable retrieval pipelinePrototype-grade security choices
Audio output as part of the workflowNot a fully generalized multi-agent system

There is one more trade-off worth naming. The repo uses patterns that are perfectly understandable in a solo project, but risky in production, including dangerous deserialization flags and tight coupling to local files. That does not weaken the article. It clarifies what this project is for.

What this project says about the RAG era

A lot of current AI software tries to win by adding more autonomy. This repo argues for the opposite move. Build smaller, clearer routes between tasks. Keep the system legible. Let retrieval do retrieval, let plotting do plotting, and let the user understand what happened.

That is why the project is interesting even if it is not the most polished tool in the category. It compresses a big set of modern AI ideas into a form you can reason about. That is a better lesson than a bigger agent.