OpenReview: The AI Code Reviewer That Proves Its Own Fixes
By marrying Claude 3.7 with ephemeral sandboxes and durable workflows, Vercel-Labs is moving beyond LLM "guessing" toward verified software engineering.
- OpenReview uses ephemeral Vercel Sandboxes to verify code fixes by running actual test suites before posting suggestions.
- Vercel Workflows prevent the agent from timing out by orchestrating long-running review tasks across stateful execution steps.
- A Just-In-Time Skill Discovery system manages context by only loading relevant documentation when the agent specifically requests it.
- The platform shifts AI code review from probabilistic text generation to a verified engineering process.
Beyond the Hallucinated Suggestion
The core problem with current AI pull request bots is fundamental. They suggest code based on statistical patterns rather than execution context. An LLM can spot a missing dependency import or a syntax error, but it cannot guarantee the proposed fix will actually compile. This leads to confident hallucinations that break the build and waste developer time.
OpenReview attacks this problem with a physical environment. Instead of relying solely on better prompting, it gives the AI a sandbox. It transforms the AI from a reviewer that tells you what is wrong into a mechanic that runs the engine, hears the knock, and proves the fix works before suggesting it.
PR reviews are a bottleneck in almost every team I’ve seen. You push a change, wait for a human reviewer to find time, get nit-picked on formatting issues that a linter should have caught, iterate, wait again. The stuff that actually matters — architecture decisions, security implications, edge cases — gets less attention because reviewers are spending mental energy on the mechanical stuff first.
A Brain with a Body (The Sandbox)
The magic happens inside workflow/steps/create-sandbox.ts. When triggered, OpenReview provisions a Vercel Sandbox. This is an ephemeral micro-VM where the bot clones the specific PR branch and installs dependencies. It does not just read the diff. It enters a live development environment.
Equipped with a bash tool, the Claude 3.7 agent can run the project's own test suite or linters. If it sees a type error, it writes a fix and runs npm test again. It only posts a GitHub suggestion block once the sandbox confirms the code passes. This is a "Code Interpreter" pattern applied directly to the pull request lifecycle.
Durable Intent: Why It Doesn't Time Out
Running a full test suite inside a sandbox takes time. Standard serverless functions time out after a few seconds or minutes, killing the AI mid-thought. OpenReview solves this by orchestrating the entire process through Vercel Workflows.
Defined in workflow/index.ts, the bot workflow manages stateful execution across multiple steps. The system can sleep while a long-running test executes and wake up exactly where it left off to process the results. If a step fails due to a network blip, the workflow resumes gracefully without restarting the entire review.
The "Just-In-Time" Expert
Managing the context window is a constant struggle for AI agents. Stuffing a prompt with every coding standard and framework best practice is expensive and degrades performance. OpenReview introduces a clever Skill Discovery system in the .agents/skills directory.
Rather than loading all knowledge upfront, the agent is presented with a menu of available skills (like Next.js routing rules or SQL security practices). When reviewing a file, the agent uses a loadSkill tool to fetch only the relevant markdown documentation. It acts like a human developer consulting the docs exactly when needed.
The New Standard for PRs
OpenReview marks a turning point from static analysis wrappers to autonomous software mechanics. By combining durable execution with isolated sandboxes, it treats code review as an engineering task rather than a text generation problem.
| Feature | Static AI Reviewers | OpenReview |
|---|---|---|
| Context Awareness | Text-based Git diffs | Full repository clone |
| Verification | LLM confidence | Live execution (npm test) |
| Execution Model | Standard Serverless (Timeout-prone) | Durable Workflows (Resumable) |
| Knowledge Loading | Bloated system prompts | Just-In-Time Skill Discovery |
OpenReview is currently in beta. It was built as an internal project to help the Vercel team test their technologies together. Expect rough edges and breaking changes.
Sources: vercel-labs/openreview Repository, Official Documentation, Towards AI coverage.