dev-browser: The Sandbox That Gives Claude a Pair of Hands

Moving beyond the "one-click-at-a-time" bottleneck with a secure, scriptable bridge between LLMs and the live web.

SawyerHood/dev-browser

A wide shot of a small figure frantically plugging cables into a massive switchboard, representing the latency of traditional AI tool-calling.
The traditional Model Context Protocol (MCP) forces agents to make a slow round-trip for every single browser action.

Every browser automation tool I've tried fails the same way. Context windows explode. Agents wander aimlessly. Your API bill climbs while the agent spends 30 turns clicking a dropdown.

Sawyer Hood, Project Creator · Closing the Agent Loop

Key Takeaways

The Latency Tax of 'Look-Think-Act'

The current standard for giving AI agents tools is the Model Context Protocol (MCP). It works, but it's inherently chatty. If an agent wants to scrape a table, it has to look at the page, think about the next step, call a 'click' tool, wait for the network round-trip, look at the new page state, and repeat.

Portrait of Sawyer Hood

This 'latency tax' makes complex web tasks prohibitively slow and expensive. dev-browser approaches the problem differently. Instead of forcing the AI to micromanage every click, it lets the AI write a script to do the job.

A Scripting Engine in a Jail

The core innovation of dev-browser is its execution model. When Claude needs to interact with the web, it doesn't send a series of discrete commands. It writes a block of JavaScript. This script is then executed locally by a background daemon.

A 3d Interactive diagram showing the execution flow. Left node: Claude (LLM). Arrow pointing right labeled 'Sends JS Script'. Middle node: QuickJS WASM Sandbox (The Jail) **important** Use metaphors and symbols in 3d space to illustrate each block. There should be a sense of space and depth. Give each node a personality. Avoid just using 3d stacked boxes. Think about each node and how to best represent its function visually.

But running AI-generated code directly on a user's machine is a massive security risk. To solve this, dev-browser uses a QuickJS WebAssembly (WASM) sandbox. The AI's script runs inside this constrained environment, completely isolated from the host operating system.

A mechanical hand reaching through a translucent, hexagonal wall to interact with gears, illustrating the WASM sandbox boundary.
The QuickJS WASM sandbox allows the agent to interact with the DOM (the gears) without ever touching the host system.

Hardened for Autonomy

The sandbox isn't just about preventing malicious code; it's about stability. The quickjs-emscripten wrapper enforces strict memory limits and CPU timeouts. If an agent writes an infinite loop, the daemon simply kills the execution, preventing the host machine from freezing.

To actually control the browser from inside this jail, dev-browser implements a custom transport layer. The sandboxed script calls a 'fake' Playwright API. These calls are serialized, passed across the WASM boundary, and executed by the real Playwright instance running in Node.js on the host.

Efficiency by the Numbers

The difference between atomic tool-calling and local scripting is stark when comparing execution models.

FeatureStandard MCP (Playwright)dev-browser
Execution ModelAtomic Tool Calls (Look-Think-Act)Sandboxed JS Scripting
LatencyHigh (Network round-trip per action)Low (Local execution loop)
Token UsageHigh (Context window bloat)Low (Single script prompt)
SecurityRelies on host permissionsQuickJS WASM Sandbox isolation

The Claude-Native Workflow

dev-browser is designed specifically to integrate with Claude Code as a 'Skill'. By pre-approving the tool in Claude's settings, developers can give the agent persistent access to a browser session.

Coding agents have a hole in their feedback loop. they can write code, run tests, check linters, iterate for hours. But they can't see what users see. For any app with a frontend, the loop never closes.

Sawyer Hood, Project Creator · Closing the Agent Loop

This persistent state—managed via 'named pages'—means an agent can navigate to a complex web app in one turn, and then write a new script to interact with that exact DOM state in the next turn, without reloading or losing context. It's a bridge that finally gives coding agents the visual feedback loop they've been missing.