DeepSearchAgent-Demo: The Research Agent That Learns What It Still Doesn’t Know

A framework-less Python demo that turns web search into a report-writing loop: structure first, search second, reflection third, then another search if the draft still has gaps.

8 min read • View on GitHub • More from 666ghj

A wide editorial scene shows a research desk with an initial query on one sheet, a stack of report cards in the center, and a thread looping one paragraph card back toward a magnifying glass and search notes. It explains that <span class=the agent does not stop at a first answer. It keeps returning to gaps and repairs them." data-prompt="Create an editorial illustration rendered entirely in black ink on a pure white background. The image depicts a wide research desk seen from a slightly elevated angle. On the left, a single sheet with an initial query lies beside a modest stack of reference clippings. In the center, a neat set of index cards forms a report outline. On the right, one paragraph card is attached to a thin thread that loops back to a magnifying glass, a second pile of search notes, and a pencil marking a missing detail. The thread visibly suggests a repair loop rather than a straight line. The artist uses tightly packed crosshatching lines layered at different angles to build up shadow and form, with clean open areas of white for highlights. The line work has the quality of a classic metal engraving, precise and deliberate, with varied line weights where bold contour lines define shapes and finer interior lines create tonal depth. The overall style evokes vintage newspaper editorial illustrations from The Economist or Wall Street Journal. No color, no gradients, no grey fills. Only black lines on white. The background MUST be pure white #FFFFFF. No paper texture, no cream, no off-white, no noise, no grain. Perfectly clean flat white background." loading="lazy">
The agent starts with structure, then circles back when a paragraph still cannot support itself.
Key Takeaways

The part most agents skip: deciding what they still do not know

Most search agents stop when they have enough text to sound confident. This one keeps a draft open beside the browser. When a paragraph is weak, the weakness is not a bug. It is the next search prompt.

That is why DeepSearchAgent-Demo is worth reading even if you already know the category. It is a demo-sized Python project, but the code reads like an argument for transparent agents. You can see the structure, the state, the node transitions, and the fallback search that fires only after reflection says the paragraph still has a hole.

The result feels less like a chatbot and more like a disciplined researcher. It builds a report skeleton first, then proves each section one at a time.

Why the outline comes before the search

The outline is the control system. ReportStructureNode turns a query into sections before the agent touches the web, so the rest of the pipeline already knows what it is trying to fill.

That choice changes the failure mode. Instead of producing a long answer that drifts, the agent can fail in a smaller place, then repair a smaller place. The code is not trying to be clever. It is trying to stay legible while it works.

query = user_question
structure = ReportStructureNode().run(query)

for paragraph in structure.paragraphs:
    results = FirstSearchNode().run(paragraph)
    draft = FirstSummaryNode().run(results)
    gap = ReflectionNode().run(paragraph, draft)

    if gap:
        results = FirstSearchNode().run(gap.followup_query)
        draft = FirstSummaryNode().run(results)
        paragraph.update(draft)

final_report = FormattingNode().run(structure)

The core trick is not search. It is the loop that notices a paragraph is still underspecified and sends it back for another pass.

Search, summarize, reflect, repair

FirstSearchNode fetches material. FirstSummaryNode turns that material into a paragraph draft. ReflectionNode compares the draft with the paragraph goal, then writes a new query if the answer is still thin.

That reflection step is the secret sauce. It does not ask, "Did I find anything?" It asks, "Did I answer the thing I said I would answer?" That shift makes the agent feel closer to an editor than a summarizer.

The state object matters because it remembers the trail. Each paragraph keeps its own research history, so the next pass can avoid redundancy and aim at the missing piece instead of spraying the same query again.

The broader DeepSearch ecosystem pushes in the same direction, toward code-first agents that make actions explicit. One contributor describes that philosophy this way:

The DeepSearchAgent project embodies the philosophy that executable code as action is the most powerful paradigm for AI agents.

lwyBZss8924d, Contributor · lwyBZss8924d/DeepSearchAgents

That is the lens here too. Code is not just implementation detail. It is the explanation.

Why the code stays readable without a framework

The architecture is small on purpose. src/agent.py orchestrates the flow, src/nodes/ keeps the steps separate, src/llms/ hides provider differences, src/state/ tracks memory, and src/tools/ wraps the web search call.

That layout makes side effects visible. If a paragraph changes, you know which node changed it. If the model provider changes, you know where the abstraction boundary lives. If JSON needs cleanup, the utility layer handles it without turning the agent into a maze.

DeepSearchAgent
  -> ReportStructureNode
  -> FirstSearchNode
  -> FirstSummaryNode
  -> ReflectionNode
  -> FormattingNode

State
  - query
  - report_title
  - paragraphs[]
  - research history

What this demo beats, and what it does not

Compared with LangChain-style orchestration or fuller deep research systems, this repo is narrower on purpose. It wins on inspectability, not breadth.

DimensionDeepSearchAgent-DemoFramework-heavy stacks
StateA single visible State object tracks paragraphs, searches, and revisions.State often lives inside chained abstractions.
Repair loopReflectionNode can generate a follow-up query from the paragraph gap.Reasoning often collapses into one long tool loop.
DebuggingEvery step is a named node in src/nodes/.Behavior can be spread across agents, chains, and callbacks.
Best useLearning, prototyping, and transparent research flows.Broader feature sets and faster assembly of complex apps.
A split editorial scene contrasts an opaque tangle of stacked blocks and arrows on the left with a clean, labeled pipeline on the right. The image explains why the repo's clarity matters: the flow is easy to trace, debug, and reason about.
The repo trades maximal automation for a pipeline you can actually inspect.

That split matters because many agent demos are impressive until you ask how a sentence was formed. Here, the answer is usually visible in the node name.

The limits are part of the lesson

This is demo-scale by design. It runs sequentially, keeps state in memory, and does not try to parallelize paragraph research or persist a long-term project database.

Those constraints are not missing features so much as boundaries. They keep the lesson sharp: control the shape of the work first, then think about scale. If you cannot trace a paragraph from query to repair, more automation will only hide the problem faster.