react-bench: React Bench: The Benchmark That Tests Whether AI Can Find the Right React File

Aiden Ybai built a deliberately cursed Next.js app to measure source-file retrieval, not code generation. The result is a benchmark for the browser era of AI coding.

8-10 min read · aidenybai/react-bench

A developer sits before a sprawling maze made of React component paths, nested panels, and misleading branches. One tiny file icon glows at the center while the rest of the maze loops around portals, fragments, and wrappers, showing how a visible UI can hide the real source file behind it.
React Bench treats source-file retrieval like a maze problem. The benchmark asks whether an agent can walk from what it sees in the browser to the one file that actually matters.

Evaluating coding agents on React retrieval tasks in complex, real-world codebases.

Aiden Ybai, Project Creator · React Bench
Key Takeaways

Why finding the file is harder than writing the code

AI coding agents are getting better at producing UI code. They are still shaky at the messier job that comes right after: mapping a rendered element back to the file where it lives. That gap is the point of aidenybai/react-bench.

The benchmark is built around a simple but uncomfortable truth. In modern React apps, the visible component tree is often a bad map of the source tree. HOCs, portals, fragments, dynamic imports, and naming tricks all weaken the old habit of searching by string and hoping the answer falls out.

What React Bench is actually measuring

React Bench is not a generic coding score. It measures source-file retrieval. A browser renders a UI. A prompt describes what the agent is looking at. Then a resolver tries to answer the only question that matters: which file produced this element?

The benchmark is a translation loop. It turns a UI observation into a source-path guess, then grades whether the guess lands on the right file fast enough.

That setup matters because it isolates a hard subproblem. Plenty of tools can tell you what component you clicked. Fewer can tell you exactly where that component came from in the codebase. React Bench is asking whether the tool helps the agent skip search, not just decorate it.

The cursed component library is the test

The repo’s most interesting move is not the dashboard. It is the gallery of intentionally ugly patterns buried in the benchmark app. The codebase includes 14-layer wrapper stacks, portal jumps, nested fragments, memo shells, hashed test IDs, and components hidden inside config objects.

A single React component sits buried at the center of layered translucent shells, each one wrapping the next like nested dolls. The outer shells suggest HOC, memo, fragment, portal, and dynamic wrapper layers, while a narrow beam points to the hidden inner component at the core, explaining how wrappers obscure source locality.
The benchmark does not invent exotic failures. It concentrates patterns that already exist in real production code, then uses them to break shallow search habits.

That is the benchmark’s bet. If a model or tool only understands the obvious structure of a React app, it will stumble here. If it can trace runtime shape back to source locality, it has a chance.

How the harness scores tools

The harness wraps every resolver in the same evaluation loop. It records the file path and component name a tool returns, compares that result with ground truth, and then combines accuracy with speed. Wrong answers are not free. They can incur a timeout-style penalty that keeps slow guesses from looking better than they are.

That choice is important. A benchmark that only counts correctness can reward bloated search. A benchmark that only counts speed can reward shallow confidence. React Bench uses the geometric mean to keep one extreme case from dominating the result, which makes the score harder to game and easier to trust.

ToolWhat it returnsDoes it help skip search?Result in React Bench
Baseline Claude CodePlain agent reasoningNoLower accuracy and slower search
Click to ComponentComponent contextNot enoughNo meaningful lift
LocatorJSComponent contextNot enoughNo meaningful lift
React GrabFile path and line numberYesTop-tier accuracy and speed
InstrucktFile path and line numberYesTop-tier accuracy and speed
AgentationFile path and line numberYesTop-tier accuracy, slower than the fastest tools
Cursor BrowserFile path and line numberYesTop-tier accuracy, slower than the fastest tools

React Bench vs the tools it tests

The comparison is not subtle. Baseline Claude Code can get far on its own, but the tools that matter are the ones that hand it a direct route to source locality. In the benchmark’s own reported runs, the simple context tools do not move accuracy much. The stronger resolvers do, because they return the one thing the agent actually needs: the file path, often plus a line number.

Without any tool, Claude Code gets 86% accuracy at a geometric mean of 45.1 seconds.

Adding React Grab, Agentation, Cursor Browser, or Instruckt pushes accuracy to 95-96%.

That is the strategic takeaway. The benchmark is less about proving that AI can read React and more about proving that the right browser-side context can collapse the search phase. The best tools do not merely identify a component. They point the agent to the exact file fast enough that the rest of the task becomes normal coding.

Why this benchmark exists

Aiden Ybai says he built this after a previous benchmark was too tidy. That makes sense. Clean dashboards are not where agents get lost. Real codebases are where they get lost, especially when the UI layer has grown a thicket of wrappers, portals, and indirection.

React Bench is a counter-design. It looks at the ugliest parts of React production code and turns them into a measured challenge. That makes it useful not just for tool builders, but for anyone betting on the browser as the next frontier for AI-assisted development.

Sources

GitHub avatar of Aiden Ybai, the creator of React Bench. The portrait identifies the project author and helps connect the benchmark to its builder.