DeepResearch-Eval: Teaching Benchmarks to Read Like Editors

HKUDS’s evaluation framework scores long AI research reports for structure, insight, redundancy, and factual grounding by pairing LLM judges with live web verification.

9 min read • View on GitHub • More from HKUDS

A wide editorial scene shows a long research report spread across a clean desk, with a magnifying glass over one column of prose and another over a stack of citations. The image explains that this project evaluates finished reports as both writing and evidence, not as a single answer string.
DeepResearch-Eval treats the report itself as the product, then asks whether it reads well and whether it holds up under source checking.
Key Takeaways

The report is the product

DeepResearch-Eval starts from a simple refusal: if an agent writes a 1,000-word research report, you should not grade it like a trivia answer. You should read it like an editor and verify it like a fact-checker. That is the repo's central move, and it is why the project feels more like publishing infrastructure than a typical benchmark.

Reports are the most canonical and representative outputs of DeepResearch. High-quality research reports feature clear structure, rigorous logic, dense information, and trustworthy citations—crucial for knowledge-intensive research scenarios.

HKUDS/DeepResearch-Eval README, Project Documentation · HKUDS/DeepResearch-Eval README

The repo ships a curated dataset of 100 queries and 100 corresponding reports generated by Qwen-DeepResearch, which gives the framework something concrete to judge. That matters, because the target is not a toy response. It is the kind of long-form output that can look polished, feel complete, and still hide weak structure or shaky sourcing.

DeepResearch-Eval splits judgment into two jobs

The architecture is clean for a research repo. judge_score.py handles report quality, while judge_fact.py handles factual verification. Underneath them, Atools.py deals with scraping and LLM plumbing, and Aprompts.py holds the rubrics that tell the judge what good looks like.

The core design is a split-brain evaluator: one path reads the report as writing, the other audits it as evidence.

The redundancy trap

The smartest part of the quality judge is not the rubric itself. It is the redundancy check. Instead of only asking whether a report covers the right topics, DeepResearch-Eval samples paragraph pairs and asks whether they are meaningfully different. That catches a failure mode most benchmarks ignore: the model that keeps talking just to keep talking.

A close-up engraving shows two nearly identical columns of text being measured with a ruler and marked with pencil ticks. One paragraph loops back on itself in a small circular path, illustrating how repetitive writing can hide inside a long report.
The project looks for repeated structure, not just wrong facts. That is how it spots padded reports that sound productive while adding little new information.

That is an editorial insight disguised as code. Good editing is partly subtraction. DeepResearch-Eval bakes that instinct into the judge by asking whether adjacent chunks of the report are helping the reader, or merely filling space.

How the fact checker decides what survives

The factuality path is more literal. It extracts claims from the report, normalizes URLs so markdown noise does not confuse the judge, scrapes live pages with Firecrawl or Jina Reader, then asks an LLM to compare the claim with the fetched context. Each claim gets a ternary verdict, which is more honest than a binary true or false.

if support == "supported":
    score = 1
elif support == "uncertain":
    score = 0
else:
    score = -1

return score, evidence

That ternary scale matters because web evidence is often messy. A claim can be partly right, loosely phrased, or impossible to verify from the available source. DeepResearch-Eval leaves room for uncertainty instead of forcing a false sense of precision.

Why prompt tables matter here

Aprompts.py is not just configuration. It is the evaluator's operating system. The quality rubric turns a fuzzy concept like insightfulness into a structured table, which is exactly what you want if you are trying to make LLM judging repeatable instead of vibes-based.

We release a carefully curated dataset: 100 queries spanning diverse categories and 100 corresponding reports generated by Qwen-DeepResearch to support systematic evaluation.

HKUDS/DeepResearch-Eval README, Project Documentation · HKUDS/DeepResearch-Eval README

This is where the repo stops feeling like a script bundle and starts feeling like a methodology. The prompts define the benchmark's standards, the judge scripts operationalize them, and the data gives those standards something real to test. In other words, the code is only half the story. The rubric is the other half.

How it compares with other deep research systems

ProjectWhat it evaluatesPrimary unitWeb verificationRedundancy check
DeepResearch-EvalFinished research reportsOutputYes, against scraped sourcesYes, paragraph pairs
OpenDeepResearchResearch generation workflowsProcessNot the main focusNo
Infinity-AILab DeepResearchEvalTask construction and agentic evaluationPipelineYes, active fact-checkingNot central
Classic QA benchmarksShort answersAnswerNoNo

The comparison makes the niche obvious. DeepResearch-Eval is not trying to build the researcher. It is trying to judge the report that comes out the other side. That puts it closer to editorial review and fact-checking than to agent orchestration.

Why this matters beyond one repo

As AI systems produce longer deliverables, evaluation has to mature with them. A report can be fluent and still be padded. It can be sourced and still be shallow. It can be correct in pieces and still fail as a whole. DeepResearch-Eval is useful because it treats those as separate problems, then gives each problem a tool that can actually see it.

That is the bigger lesson here. Once models write work products instead of short answers, benchmarks need to look less like multiple-choice tests and more like editorial review boards. This repo is a compact version of that future.