OpenReview: The AI Code Reviewer That Proves Its Own Fixes

By marrying Claude 3.7 with ephemeral sandboxes and durable workflows, Vercel-Labs is moving beyond LLM "guessing" toward verified software engineering.

8 min read • View on GitHub • More from vercel-labs

A vintage-style laboratory scale balancing a code snippet with a heavy iron checkmark, illustrating the concept of verified code suggestions.
OpenReview shifts AI from a passive commentator to an active mechanic that verifies its own work.
Key Takeaways

Beyond the Hallucinated Suggestion

The core problem with current AI pull request bots is fundamental. They suggest code based on statistical patterns rather than execution context. An LLM can spot a missing dependency import or a syntax error, but it cannot guarantee the proposed fix will actually compile. This leads to confident hallucinations that break the build and waste developer time.

OpenReview attacks this problem with a physical environment. Instead of relying solely on better prompting, it gives the AI a sandbox. It transforms the AI from a reviewer that tells you what is wrong into a mechanic that runs the engine, hears the knock, and proves the fix works before suggesting it.

PR reviews are a bottleneck in almost every team I’ve seen. You push a change, wait for a human reviewer to find time, get nit-picked on formatting issues that a linter should have caught, iterate, wait again. The stuff that actually matters — architecture decisions, security implications, edge cases — gets less attention because reviewers are spending mental energy on the mechanical stuff first.

— Gowtham Boyina, Towards AI

A Brain with a Body (The Sandbox)

The magic happens inside workflow/steps/create-sandbox.ts. When triggered, OpenReview provisions a Vercel Sandbox. This is an ephemeral micro-VM where the bot clones the specific PR branch and installs dependencies. It does not just read the diff. It enters a live development environment.

Equipped with a bash tool, the Claude 3.7 agent can run the project's own test suite or linters. If it sees a type error, it writes a fix and runs npm test again. It only posts a GitHub suggestion block once the sandbox confirms the code passes. This is a "Code Interpreter" pattern applied directly to the pull request lifecycle.

A flowchart showing the Verified Loop. The sequence starts at a 'GitHub PR' node

Durable Intent: Why It Doesn't Time Out

Running a full test suite inside a sandbox takes time. Standard serverless functions time out after a few seconds or minutes, killing the AI mid-thought. OpenReview solves this by orchestrating the entire process through Vercel Workflows.

Defined in workflow/index.ts, the bot workflow manages stateful execution across multiple steps. The system can sleep while a long-running test executes and wake up exactly where it left off to process the results. If a step fails due to a network blip, the workflow resumes gracefully without restarting the entire review.

The "Just-In-Time" Expert

Managing the context window is a constant struggle for AI agents. Stuffing a prompt with every coding standard and framework best practice is expensive and degrades performance. OpenReview introduces a clever Skill Discovery system in the .agents/skills directory.

Rather than loading all knowledge upfront, the agent is presented with a menu of available skills (like Next.js routing rules or SQL security practices). When reviewing a file, the agent uses a loadSkill tool to fetch only the relevant markdown documentation. It acts like a human developer consulting the docs exactly when needed.

A mechanical hand pulling a single specific folder labeled 'LINT RULES' from a massive filing cabinet.
Skill Discovery allows the agent to load domain-specific knowledge only when context dictates it is necessary.

The New Standard for PRs

OpenReview marks a turning point from static analysis wrappers to autonomous software mechanics. By combining durable execution with isolated sandboxes, it treats code review as an engineering task rather than a text generation problem.

Feature Static AI Reviewers OpenReview
Context Awareness Text-based Git diffs Full repository clone
Verification LLM confidence Live execution (npm test)
Execution Model Standard Serverless (Timeout-prone) Durable Workflows (Resumable)
Knowledge Loading Bloated system prompts Just-In-Time Skill Discovery

OpenReview is currently in beta. It was built as an internal project to help the Vercel team test their technologies together. Expect rough edges and breaking changes.

— OpenReview Documentation, Project Documentation

Sources: vercel-labs/openreview Repository, Official Documentation, Towards AI coverage.