react-bench: React Bench: The Benchmark That Tests Whether AI Can Find the Right React File
Aiden Ybai built a deliberately cursed Next.js app to measure source-file retrieval, not code generation. The result is a benchmark for the browser era of AI coding.
Evaluating coding agents on React retrieval tasks in complex, real-world codebases.
- React Bench argues that AI coding pain begins after code generation, when the model has to locate the right file behind a rendered UI.
- Its cursed component gallery turns real React anti-patterns into a controlled test for source-file retrieval.
- The benchmark rewards tools that return file path and line number, not just component names or clickable overlays.
- The real comparison is between search and skip-search: which tools help an agent jump straight to source locality.
Why finding the file is harder than writing the code
AI coding agents are getting better at producing UI code. They are still shaky at the messier job that comes right after: mapping a rendered element back to the file where it lives. That gap is the point of aidenybai/react-bench.
The benchmark is built around a simple but uncomfortable truth. In modern React apps, the visible component tree is often a bad map of the source tree. HOCs, portals, fragments, dynamic imports, and naming tricks all weaken the old habit of searching by string and hoping the answer falls out.
What React Bench is actually measuring
React Bench is not a generic coding score. It measures source-file retrieval. A browser renders a UI. A prompt describes what the agent is looking at. Then a resolver tries to answer the only question that matters: which file produced this element?
That setup matters because it isolates a hard subproblem. Plenty of tools can tell you what component you clicked. Fewer can tell you exactly where that component came from in the codebase. React Bench is asking whether the tool helps the agent skip search, not just decorate it.
The cursed component library is the test
The repo’s most interesting move is not the dashboard. It is the gallery of intentionally ugly patterns buried in the benchmark app. The codebase includes 14-layer wrapper stacks, portal jumps, nested fragments, memo shells, hashed test IDs, and components hidden inside config objects.
That is the benchmark’s bet. If a model or tool only understands the obvious structure of a React app, it will stumble here. If it can trace runtime shape back to source locality, it has a chance.
How the harness scores tools
The harness wraps every resolver in the same evaluation loop. It records the file path and component name a tool returns, compares that result with ground truth, and then combines accuracy with speed. Wrong answers are not free. They can incur a timeout-style penalty that keeps slow guesses from looking better than they are.
That choice is important. A benchmark that only counts correctness can reward bloated search. A benchmark that only counts speed can reward shallow confidence. React Bench uses the geometric mean to keep one extreme case from dominating the result, which makes the score harder to game and easier to trust.
| Tool | What it returns | Does it help skip search? | Result in React Bench |
|---|---|---|---|
| Baseline Claude Code | Plain agent reasoning | No | Lower accuracy and slower search |
| Click to Component | Component context | Not enough | No meaningful lift |
| LocatorJS | Component context | Not enough | No meaningful lift |
| React Grab | File path and line number | Yes | Top-tier accuracy and speed |
| Instruckt | File path and line number | Yes | Top-tier accuracy and speed |
| Agentation | File path and line number | Yes | Top-tier accuracy, slower than the fastest tools |
| Cursor Browser | File path and line number | Yes | Top-tier accuracy, slower than the fastest tools |
React Bench vs the tools it tests
The comparison is not subtle. Baseline Claude Code can get far on its own, but the tools that matter are the ones that hand it a direct route to source locality. In the benchmark’s own reported runs, the simple context tools do not move accuracy much. The stronger resolvers do, because they return the one thing the agent actually needs: the file path, often plus a line number.
Without any tool, Claude Code gets 86% accuracy at a geometric mean of 45.1 seconds.
Adding React Grab, Agentation, Cursor Browser, or Instruckt pushes accuracy to 95-96%.
That is the strategic takeaway. The benchmark is less about proving that AI can read React and more about proving that the right browser-side context can collapse the search phase. The best tools do not merely identify a component. They point the agent to the exact file fast enough that the rest of the task becomes normal coding.
Why this benchmark exists
Aiden Ybai says he built this after a previous benchmark was too tidy. That makes sense. Clean dashboards are not where agents get lost. Real codebases are where they get lost, especially when the UI layer has grown a thicket of wrappers, portals, and indirection.
React Bench is a counter-design. It looks at the ugliest parts of React production code and turns them into a measured challenge. That makes it useful not just for tool builders, but for anyone betting on the browser as the next frontier for AI-assisted development.
Sources