Document-Retrieval: The YouTube Chatbot That Fills Its Own Knowledge Gap

A LangGraph agent, BrightData scraper, and pgvector-backed pipeline turn missing context into a fetchable task instead of a dead end.

7 to 9 min read • View on GitHub • More from riyav1606

A wide editorial scene shows a desktop question terminal branching into two paths. One arm checks a cabinet of existing transcripts while another extends through a window toward a YouTube screen and a scraping rig, then feeds pages into an index box. The scene explains the repo's core idea: if the answer is missing, the system can go acquire the source first.
The thesis is the loop, not the stack. This repo can decide to fetch the missing transcript before it tries to answer.
Key Takeaways

Most RAG apps search what they already have. This one can go get what it needs first. That is the difference between a chatbot with a vector store and a chatbot that behaves like a researcher.

The chatbot that calls in the source material

The repo is built around a simple but sharp idea: if a question lands outside the current index, the system should not stall. It should decide whether to search known material, look across related videos, or trigger a fresh scrape and make the missing transcript searchable.

A simple document retrieval system built with Python and Flask.

Riya, Project Creator/Maintainer · Document-Retrieval

That README line is modest, but the behavior underneath is more ambitious. It turns retrieval into a decision loop, which is a better fit for content that is not always preloaded.

How the retrieval loop closes itself

The key move is the asynchronous handoff. The agent can keep the conversation moving while a webhook finishes the scrape and updates the index.

The architecture matters because it avoids a dead end. The agent receives a question, chooses a tool, and only then decides whether it needs to fetch more material. If a scrape is required, the request leaves the conversational loop, runs externally, and comes back through a webhook once the transcript is ready.

A close-up relay shows a question card, a webhook parcel moving between two machines, and a stamped transcript tile landing into a grid of vector nodes. The image explains how the scraper can finish later without freezing the conversation.
The webhook is the bridge. It lets a long-running scrape become a normal step in a conversational workflow.

Inside the agent: tool choice is the product

The agent is built with LangGraph's createReactAgent, which matters because the workflow is stateful, not just prompt-driven. The model is not asked to hallucinate a single best search step. It is given tools that map to different kinds of uncertainty.

const tools = [
  retrieveTool,
  retrieveSimilarVideosTool,
  triggerYoutubeVideoScrapeTool,
];

const agent = createReactAgent({
  llm: claude,
  tools,
  checkpointer: new MemorySaver(),
});

That split is the product. One tool answers questions about a known video. Another searches across videos. The third creates new searchable material when the corpus is too thin to answer responsibly.

ToolJobWhy it matters
retrieveToolSearch a specific video's transcriptBest when the user already knows the video target
retrieveSimilarVideosToolSearch across videos by semantic intentUseful when the user knows the topic but not the exact source
triggerYoutubeVideoScrapeToolStart a fresh scrape and indexing passTurns missing context into a solvable acquisition step

Why the pipeline needs a webhook

BrightData scraping is not a quick local function call. It is an external job with enough latency that blocking the chat would make the app feel broken. The webhook pattern solves that cleanly: start the scrape, return control, then re-enter the system when the transcript is ready.

That design also matches the user's mental model. The bot can say, in effect, "I am fetching that source now," instead of pretending it already knows the answer. The conversation stays coherent even while the data pipeline is still working.

Why the vector store is simple, and that is the point

The retrieval layer uses familiar pieces: recursive character splitting, embeddings, and pgvector. The chunk size is 1000 with 200 characters of overlap, which is a conservative choice that keeps boundaries from cutting too aggressively through transcript context.

const splitter = new RecursiveCharacterTextSplitter({
  chunkSize: 1000,
  chunkOverlap: 200,
});

const vectorStore = await PGVectorStore.initialize(embeddings, {
  postgresConnectionOptions,
});

Nothing here is flashy. That is the virtue. The repo is not trying to invent a new retrieval primitive. It is using standard parts well enough that the interesting behavior can happen above them.

ApproachCorpus assumptionMissing knowledge behaviorLatency modelBest fitLimitation
Document-RetrievalCorpus can be acquired on demandTriggers a scrape and re-indexesAsync, webhook-backedConversational research over videosPrototype-level polish
Static RAGCorpus must already existFails or guessesFast, but bounded by preloadStable internal knowledge basesNo path to new sources
Traditional search stackCorpus is managed externallySearches only what is indexedOptimized for retrieval, not acquisitionLarge-scale site searchNo agentic fetch loop

A practical prototype, not a finished product

The repo feels like an exploration, not a polished product. There is a hardcoded transcript for testing, no README to explain the story, and the frontend is functional rather than ornate. That does not weaken the idea. It makes the architecture easier to inspect.

This is a strong prototype because it exposes the right pattern early. The project is less about a YouTube chatbot and more about a system that can create its own retrieval target when necessary.

What this repo teaches about AI apps

The larger lesson is that good AI products will not just answer from what they already know. They will know when to gather more evidence, how to make that evidence searchable, and how to resume the conversation once the pipeline catches up.

PatternWhat it optimizes forWhat this repo adds
Search-only chatLow latency answersNothing beyond the existing index
Agentic retrievalChoosing the right sourceA path to fetch missing context
Acquisition-first AIAnswer quality under uncertaintyThe ability to build the corpus mid-conversation

In that sense, Document-Retrieval is less a search engine demo than a product pattern demo. It shows how a small system can move from query to acquisition to indexed evidence without making the user restart the task.