DeepSearchAgent-Demo: The Research Agent That Learns What It Still Doesn’t Know
A framework-less Python demo that turns web search into a report-writing loop: structure first, search second, reflection third, then another search if the draft still has gaps.
the agent does not stop at a first answer. It keeps returning to gaps and repairs them." data-prompt="Create an editorial illustration rendered entirely in black ink on a pure white background. The image depicts a wide research desk seen from a slightly elevated angle. On the left, a single sheet with an initial query lies beside a modest stack of reference clippings. In the center, a neat set of index cards forms a report outline. On the right, one paragraph card is attached to a thin thread that loops back to a magnifying glass, a second pile of search notes, and a pencil marking a missing detail. The thread visibly suggests a repair loop rather than a straight line. The artist uses tightly packed crosshatching lines layered at different angles to build up shadow and form, with clean open areas of white for highlights. The line work has the quality of a classic metal engraving, precise and deliberate, with varied line weights where bold contour lines define shapes and finer interior lines create tonal depth. The overall style evokes vintage newspaper editorial illustrations from The Economist or Wall Street Journal. No color, no gradients, no grey fills. Only black lines on white. The background MUST be pure white #FFFFFF. No paper texture, no cream, no off-white, no noise, no grain. Perfectly clean flat white background." loading="lazy">
- DeepSearchAgent-Demo treats research as a repair loop, not a one-shot summary, so the draft improves when the agent notices what it cannot yet prove.
- The outline comes first, which keeps the system from wandering and gives every paragraph its own job.
- Its value is readability: named nodes, visible state, and a small abstraction layer make the workflow easy to inspect.
- That clarity comes with trade-offs, because the demo stays sequential, in-memory, and intentionally small.
The part most agents skip: deciding what they still do not know
Most search agents stop when they have enough text to sound confident. This one keeps a draft open beside the browser. When a paragraph is weak, the weakness is not a bug. It is the next search prompt.
That is why DeepSearchAgent-Demo is worth reading even if you already know the category. It is a demo-sized Python project, but the code reads like an argument for transparent agents. You can see the structure, the state, the node transitions, and the fallback search that fires only after reflection says the paragraph still has a hole.
The result feels less like a chatbot and more like a disciplined researcher. It builds a report skeleton first, then proves each section one at a time.
Why the outline comes before the search
The outline is the control system. ReportStructureNode turns a query into sections before the agent touches the web, so the rest of the pipeline already knows what it is trying to fill.
That choice changes the failure mode. Instead of producing a long answer that drifts, the agent can fail in a smaller place, then repair a smaller place. The code is not trying to be clever. It is trying to stay legible while it works.
query = user_question
structure = ReportStructureNode().run(query)
for paragraph in structure.paragraphs:
results = FirstSearchNode().run(paragraph)
draft = FirstSummaryNode().run(results)
gap = ReflectionNode().run(paragraph, draft)
if gap:
results = FirstSearchNode().run(gap.followup_query)
draft = FirstSummaryNode().run(results)
paragraph.update(draft)
final_report = FormattingNode().run(structure)
Search, summarize, reflect, repair
FirstSearchNode fetches material. FirstSummaryNode turns that material into a paragraph draft. ReflectionNode compares the draft with the paragraph goal, then writes a new query if the answer is still thin.
That reflection step is the secret sauce. It does not ask, "Did I find anything?" It asks, "Did I answer the thing I said I would answer?" That shift makes the agent feel closer to an editor than a summarizer.
The state object matters because it remembers the trail. Each paragraph keeps its own research history, so the next pass can avoid redundancy and aim at the missing piece instead of spraying the same query again.
The broader DeepSearch ecosystem pushes in the same direction, toward code-first agents that make actions explicit. One contributor describes that philosophy this way:
The DeepSearchAgent project embodies the philosophy that executable code as action is the most powerful paradigm for AI agents.
That is the lens here too. Code is not just implementation detail. It is the explanation.
Why the code stays readable without a framework
The architecture is small on purpose. src/agent.py orchestrates the flow, src/nodes/ keeps the steps separate, src/llms/ hides provider differences, src/state/ tracks memory, and src/tools/ wraps the web search call.
That layout makes side effects visible. If a paragraph changes, you know which node changed it. If the model provider changes, you know where the abstraction boundary lives. If JSON needs cleanup, the utility layer handles it without turning the agent into a maze.
DeepSearchAgent
-> ReportStructureNode
-> FirstSearchNode
-> FirstSummaryNode
-> ReflectionNode
-> FormattingNode
State
- query
- report_title
- paragraphs[]
- research history
What this demo beats, and what it does not
Compared with LangChain-style orchestration or fuller deep research systems, this repo is narrower on purpose. It wins on inspectability, not breadth.
| Dimension | DeepSearchAgent-Demo | Framework-heavy stacks |
|---|---|---|
| State | A single visible State object tracks paragraphs, searches, and revisions. | State often lives inside chained abstractions. |
| Repair loop | ReflectionNode can generate a follow-up query from the paragraph gap. | Reasoning often collapses into one long tool loop. |
| Debugging | Every step is a named node in src/nodes/. | Behavior can be spread across agents, chains, and callbacks. |
| Best use | Learning, prototyping, and transparent research flows. | Broader feature sets and faster assembly of complex apps. |
That split matters because many agent demos are impressive until you ask how a sentence was formed. Here, the answer is usually visible in the node name.
The limits are part of the lesson
This is demo-scale by design. It runs sequentially, keeps state in memory, and does not try to parallelize paragraph research or persist a long-term project database.
Those constraints are not missing features so much as boundaries. They keep the lesson sharp: control the shape of the work first, then think about scale. If you cannot trace a paragraph from query to repair, more automation will only hide the problem faster.