DeepResearch-Eval: Teaching Benchmarks to Read Like Editors
HKUDS’s evaluation framework scores long AI research reports for structure, insight, redundancy, and factual grounding by pairing LLM judges with live web verification.
- DeepResearch-Eval treats long research reports as finished editorial artifacts, not as answer strings.
- Its sharpest move is to split evaluation into writing quality and factual grounding, because those failures do not always travel together.
- The redundancy check matters because AI reports often fail by padding, looping, or restating the same thought in new words.
- The rubric tables turn subjective review into a repeatable judge that can score structure, insight, and support with less hand-waving.
The report is the product
DeepResearch-Eval starts from a simple refusal: if an agent writes a 1,000-word research report, you should not grade it like a trivia answer. You should read it like an editor and verify it like a fact-checker. That is the repo's central move, and it is why the project feels more like publishing infrastructure than a typical benchmark.
Reports are the most canonical and representative outputs of DeepResearch. High-quality research reports feature clear structure, rigorous logic, dense information, and trustworthy citations—crucial for knowledge-intensive research scenarios.
The repo ships a curated dataset of 100 queries and 100 corresponding reports generated by Qwen-DeepResearch, which gives the framework something concrete to judge. That matters, because the target is not a toy response. It is the kind of long-form output that can look polished, feel complete, and still hide weak structure or shaky sourcing.
DeepResearch-Eval splits judgment into two jobs
The architecture is clean for a research repo. judge_score.py handles report quality, while judge_fact.py handles factual verification. Underneath them, Atools.py deals with scraping and LLM plumbing, and Aprompts.py holds the rubrics that tell the judge what good looks like.
The redundancy trap
The smartest part of the quality judge is not the rubric itself. It is the redundancy check. Instead of only asking whether a report covers the right topics, DeepResearch-Eval samples paragraph pairs and asks whether they are meaningfully different. That catches a failure mode most benchmarks ignore: the model that keeps talking just to keep talking.
That is an editorial insight disguised as code. Good editing is partly subtraction. DeepResearch-Eval bakes that instinct into the judge by asking whether adjacent chunks of the report are helping the reader, or merely filling space.
How the fact checker decides what survives
The factuality path is more literal. It extracts claims from the report, normalizes URLs so markdown noise does not confuse the judge, scrapes live pages with Firecrawl or Jina Reader, then asks an LLM to compare the claim with the fetched context. Each claim gets a ternary verdict, which is more honest than a binary true or false.
if support == "supported":
score = 1
elif support == "uncertain":
score = 0
else:
score = -1
return score, evidence
That ternary scale matters because web evidence is often messy. A claim can be partly right, loosely phrased, or impossible to verify from the available source. DeepResearch-Eval leaves room for uncertainty instead of forcing a false sense of precision.
Why prompt tables matter here
Aprompts.py is not just configuration. It is the evaluator's operating system. The quality rubric turns a fuzzy concept like insightfulness into a structured table, which is exactly what you want if you are trying to make LLM judging repeatable instead of vibes-based.
We release a carefully curated dataset: 100 queries spanning diverse categories and 100 corresponding reports generated by Qwen-DeepResearch to support systematic evaluation.
This is where the repo stops feeling like a script bundle and starts feeling like a methodology. The prompts define the benchmark's standards, the judge scripts operationalize them, and the data gives those standards something real to test. In other words, the code is only half the story. The rubric is the other half.
How it compares with other deep research systems
| Project | What it evaluates | Primary unit | Web verification | Redundancy check |
|---|---|---|---|---|
| DeepResearch-Eval | Finished research reports | Output | Yes, against scraped sources | Yes, paragraph pairs |
| OpenDeepResearch | Research generation workflows | Process | Not the main focus | No |
| Infinity-AILab DeepResearchEval | Task construction and agentic evaluation | Pipeline | Yes, active fact-checking | Not central |
| Classic QA benchmarks | Short answers | Answer | No | No |
The comparison makes the niche obvious. DeepResearch-Eval is not trying to build the researcher. It is trying to judge the report that comes out the other side. That puts it closer to editorial review and fact-checking than to agent orchestration.
Why this matters beyond one repo
As AI systems produce longer deliverables, evaluation has to mature with them. A report can be fluent and still be padded. It can be sourced and still be shallow. It can be correct in pieces and still fail as a whole. DeepResearch-Eval is useful because it treats those as separate problems, then gives each problem a tool that can actually see it.
That is the bigger lesson here. Once models write work products instead of short answers, benchmarks need to look less like multiple-choice tests and more like editorial review boards. This repo is a compact version of that future.