rishav-RG/Multi-Agent-Chat-System: The Router Is the Real Product

A deep look at a fail-safe agent stack that chooses between search, summary, or both, routes tasks across multiple models, and keeps the whole system observable and recoverable.

9 min read • View on GitHub • More from rishav-RG

A wide black-ink editorial illustration of a railroad junction with a central dispatcher deciding which track each cargo train should take. Three main tracks diverge toward search, summary, and parallel execution, while a side lever represents the fallback route. It explains that the system’s core value is routing intelligence, not a single chatbot reply.
The project behaves like a dispatch layer. It decides where a request should go before it decides what to say.
Key Takeaways

The Chatbot That Refuses to Guess

Most chat apps try to answer first and organize later. This one does the opposite. A supervisor decides whether a request is a greeting, a search task, a summary task, or a parallel job, and if the model-based route gets shaky, the system falls back to keywords instead of pretending confidence it does not have.

That matters because agentic apps fail in boring ways. They misclassify intent, pick the wrong tool, or produce a fluent answer that never touched the right data. This repo treats those failure modes as routing problems, which is a more durable way to think about the product.

A request becomes a workflow. The diagram shows how the system checks cache, classifies intent, falls back safely, and merges branches into one answer.

Why the Workflow Matters More Than the Prompt

The brain of the repo is backend/app/workflows/chat_workflow.py. It manages a shared GraphState, checks for a cache hit, and only then hands the request to the supervisor. That sequence changes the whole shape of the app. The system is not a single turn generator. It is a stateful decision graph.

In practical terms, that means the graph can branch cleanly without turning into nested conditionals everywhere. A search result can feed summary work, a parallel route can fan out and merge, and the final response can be assembled only after the system knows what kind of problem it is solving.

That is the real design move here. The workflow does not merely call agents. It turns uncertainty into structure.

A Supervisor With a Backup Plan

The supervisor agent is where the project becomes opinionated. According to the repository structure and implementation notes, it tries LLM-based routing first and then falls back to keyword logic if classification fails. That is not glamorous, but it is exactly the sort of boring resilience that makes agent systems usable.

You can read the philosophy in the control flow. The system does not demand that the model always be right. It asks the model to help, then keeps a second path ready for when inference is unstable or ambiguous.

One Session, Multiple Models

The repo’s model strategy is more interesting than a default chatbot stack. In backend/app/services/llm_service.py, the code maps task types to different open-source models through a Hugging Face router. Summarization, code generation, and question answering are not forced through one catch-all model. They are assigned where they fit best.

TaskThis repoSingle-model chat stack
SummarizationRoutes to a specialized model through the capability mapUses the same model as every other turn
Code generationCan be sent to a code-oriented modelRelies on whatever model happens to be default
Question answeringUses a model chosen for the taskAsks one model to do everything
Operational effectBetter fit, easier tuning, clearer trade-offsSimpler to start, harder to optimize
A close-up black-ink illustration of a mechanical control panel with gauges for intent routing, keyword fallback, cache hit, and parallel execution. One gauge flickers red, and a mechanical arm switches the system to a backup path while the main machinery continues running. It explains how the repo is designed to fail softly instead of collapsing.
The engineering idea is not perfection. It is graceful fallback under uncertainty.

Memory That Fails Softly

Redis is doing two jobs here: short-term memory and cache. That is a practical choice because it keeps repeated queries cheap while preserving enough conversational state to make the workflow useful. The codebase also reportedly falls back to an in-memory store when Redis is unavailable, which is exactly the kind of local-dev resilience that saves time and reduces false alarms.

This is where production shape starts to show. The system is not pretending the infrastructure is always there. It is built to keep moving when part of the stack is missing.

Why This Feels Production-Shaped

The surrounding middleware matters as much as the agent logic. Rate limiting, security headers, structured logging, and Langfuse traces turn the repo from a demo into something closer to a deployable scaffold. In agent systems, observability is not a luxury. If routing decisions are probabilistic, you need a trail.

That is the subtle strength of the project. It recognizes that a useful AI app is not just about response quality. It is about knowing what happened, why it happened, and how to recover when the decision path was wrong.

How It Compares to the Usual Suspects

The easiest way to place this repo is by contrast. LangChain and Semantic Kernel are broad frameworks. AutoGPT leans toward autonomous execution. Flowise favors low-code flow design. This project is narrower than all of them, and that is an advantage if your goal is a code-first multi-agent chat control plane with fallback routing.

ProjectPrimary shapeStrengthWhere this repo differs
LangChainGeneral composability frameworkHuge ecosystem and flexibilityThis repo is narrower and more opinionated
AutoGPTAutonomous task runnerGoal decomposition and tool useThis repo is built around chat routing, not open-ended autonomy
Semantic KernelLLM integration SDKStructured extensibilityThis repo ships with a more specific multi-agent workflow
FlowiseVisual flow builderLow-code accessibilityThis repo is code-first and easier to shape around fallback logic

The point is not that this project beats the larger ecosystems. The point is that it makes a sharper promise. It gives you a working pattern for dispatch, branch, fallback, memory, and traceability without asking you to assemble the whole philosophy yourself.

The Bigger Lesson

The most useful AI systems will not look like one giant assistant. They will look like dispatch layers with memory, branch logic, and recovery paths. That is the lesson hiding inside this repo. It treats the model as one part of the machine, not the machine itself.

That makes the project a strong reference point for developers who want practical agent design. Not more chat. Better control.