Document-Retrieval: The YouTube Chatbot That Fills Its Own Knowledge Gap
A LangGraph agent, BrightData scraper, and pgvector-backed pipeline turn missing context into a fetchable task instead of a dead end.
- Document-Retrieval is interesting because it treats missing context as a task the agent can solve, not a failure state.
- The real product move is the closed loop from question to scrape to webhook to index to answer.
- The repo uses standard retrieval parts, but the orchestration makes them behave like an on-demand research system.
- Its lesson is broader than YouTube: useful AI apps increasingly need an acquisition layer, not just a response layer.
Most RAG apps search what they already have. This one can go get what it needs first. That is the difference between a chatbot with a vector store and a chatbot that behaves like a researcher.
The chatbot that calls in the source material
The repo is built around a simple but sharp idea: if a question lands outside the current index, the system should not stall. It should decide whether to search known material, look across related videos, or trigger a fresh scrape and make the missing transcript searchable.
A simple document retrieval system built with Python and Flask.
That README line is modest, but the behavior underneath is more ambitious. It turns retrieval into a decision loop, which is a better fit for content that is not always preloaded.
How the retrieval loop closes itself
The architecture matters because it avoids a dead end. The agent receives a question, chooses a tool, and only then decides whether it needs to fetch more material. If a scrape is required, the request leaves the conversational loop, runs externally, and comes back through a webhook once the transcript is ready.
Inside the agent: tool choice is the product
The agent is built with LangGraph's createReactAgent, which matters because the workflow is stateful, not just prompt-driven. The model is not asked to hallucinate a single best search step. It is given tools that map to different kinds of uncertainty.
const tools = [
retrieveTool,
retrieveSimilarVideosTool,
triggerYoutubeVideoScrapeTool,
];
const agent = createReactAgent({
llm: claude,
tools,
checkpointer: new MemorySaver(),
});
That split is the product. One tool answers questions about a known video. Another searches across videos. The third creates new searchable material when the corpus is too thin to answer responsibly.
| Tool | Job | Why it matters |
|---|---|---|
| retrieveTool | Search a specific video's transcript | Best when the user already knows the video target |
| retrieveSimilarVideosTool | Search across videos by semantic intent | Useful when the user knows the topic but not the exact source |
| triggerYoutubeVideoScrapeTool | Start a fresh scrape and indexing pass | Turns missing context into a solvable acquisition step |
Why the pipeline needs a webhook
BrightData scraping is not a quick local function call. It is an external job with enough latency that blocking the chat would make the app feel broken. The webhook pattern solves that cleanly: start the scrape, return control, then re-enter the system when the transcript is ready.
That design also matches the user's mental model. The bot can say, in effect, "I am fetching that source now," instead of pretending it already knows the answer. The conversation stays coherent even while the data pipeline is still working.
Why the vector store is simple, and that is the point
The retrieval layer uses familiar pieces: recursive character splitting, embeddings, and pgvector. The chunk size is 1000 with 200 characters of overlap, which is a conservative choice that keeps boundaries from cutting too aggressively through transcript context.
const splitter = new RecursiveCharacterTextSplitter({
chunkSize: 1000,
chunkOverlap: 200,
});
const vectorStore = await PGVectorStore.initialize(embeddings, {
postgresConnectionOptions,
});
Nothing here is flashy. That is the virtue. The repo is not trying to invent a new retrieval primitive. It is using standard parts well enough that the interesting behavior can happen above them.
| Approach | Corpus assumption | Missing knowledge behavior | Latency model | Best fit | Limitation |
|---|---|---|---|---|---|
| Document-Retrieval | Corpus can be acquired on demand | Triggers a scrape and re-indexes | Async, webhook-backed | Conversational research over videos | Prototype-level polish |
| Static RAG | Corpus must already exist | Fails or guesses | Fast, but bounded by preload | Stable internal knowledge bases | No path to new sources |
| Traditional search stack | Corpus is managed externally | Searches only what is indexed | Optimized for retrieval, not acquisition | Large-scale site search | No agentic fetch loop |
A practical prototype, not a finished product
The repo feels like an exploration, not a polished product. There is a hardcoded transcript for testing, no README to explain the story, and the frontend is functional rather than ornate. That does not weaken the idea. It makes the architecture easier to inspect.
This is a strong prototype because it exposes the right pattern early. The project is less about a YouTube chatbot and more about a system that can create its own retrieval target when necessary.
What this repo teaches about AI apps
The larger lesson is that good AI products will not just answer from what they already know. They will know when to gather more evidence, how to make that evidence searchable, and how to resume the conversation once the pipeline catches up.
| Pattern | What it optimizes for | What this repo adds |
|---|---|---|
| Search-only chat | Low latency answers | Nothing beyond the existing index |
| Agentic retrieval | Choosing the right source | A path to fetch missing context |
| Acquisition-first AI | Answer quality under uncertainty | The ability to build the corpus mid-conversation |
In that sense, Document-Retrieval is less a search engine demo than a product pattern demo. It shows how a small system can move from query to acquisition to indexed evidence without making the user restart the task.