sandbox-benchmarking: Measuring the 6-Second Ghost in Vercel Sandbox

How a minimal benchmarking suite exposed the hidden first-command penalty haunting ephemeral compute.

• View on GitHub • More from vercel-labs

A futuristic racing car crossing a finish line labeled 'CREATED' while the driver is still frantically trying to start the engine.
The Sandbox API reports readiness before the underlying microVM has actually finished booting.

`Sandbox.create()` resolves before the VM is fully ready. The first `runCommand` absorbs the remaining boot time. Every subsequent command is fast.

johnlindquist, Contributor · vercel-labs/sandbox-benchmarking

Key Takeaways

The Invisible Wall

Cloud providers market their ephemeral compute environments with promises of sub-second startup times. For Vercel Sandbox, the marketing highlights millisecond readiness. But developers building interactive AI agents began noticing a phantom delay. The infrastructure felt fast, yet the first response always dragged.

The repository vercel-labs/sandbox-benchmarking is a technical detective story. It operates as a bug report disguised as a benchmarking suite. Developer John Lindquist built it to isolate a specific bottleneck in Vercel's Firecracker-based microVMs. The finding is stark. The system suffers from a massive first-command penalty.

The Leaky Abstraction

The core issue lies in how the Vercel Sandbox SDK handles state. When a developer calls Sandbox.create(), the API returns a success response almost immediately. The hypervisor has allocated the resources. To the control plane, the sandbox is running.

But running does not mean ready. The underlying operating system is still booting. When the developer sends the very first sandbox.runCommand(), that command hits a queue. It sits there, waiting for the VM to actually accept instructions. The SDK makes the command execution look incredibly slow, but the command is simply absorbing the hidden boot time.

How the first command absorbs the VM boot time, creating a massive artificial delay.

WSJ-style hedcut portrait of John Lindquist

Benchmarking the Benchmarks

To prove this wasn't just local network latency, the repository includes two identical test environments. One runs locally via bench.mjs. The other runs inside Vercel's own infrastructure as a serverless function via api/bench.mjs. Both scripts execute a trivial command twice in a row.

Minimal reproduction: first runCommand after Sandbox.create() has 7-112x latency penalty vs subsequent commands

johnlindquist, Contributor · vercel-labs/sandbox-benchmarking
Environmentcreate() LatencyFirst Command (Penalty)Second Command (Ready)
Local CLI~1.5s~2.0s - 4.0s< 100ms
Vercel Function (IAD1)~800ms~6.5s - 33.0s< 100ms

The data reveals a counter-intuitive truth. Running the benchmark from within Vercel's own network often results in a worse first-command penalty than running it over the public internet from a local machine.

The Snapshot Irony

The most surprising revelation involves snapshots. Snapshots are designed to be an optimization. You freeze a VM state and restore it later to skip the boot sequence. But the benchmark shows that restoring from a snapshot can actually be slower than starting fresh.

A person struggling to move a massive, overstuffed suitcase anchored to the floor, while an empty suitcase slides easily.
Restoring from a warm snapshot can incur a higher penalty than provisioning a fresh, empty environment.

The production penalty is **~6 seconds on every restore**, consistently.

johnlindquist, Contributor · vercel-labs/sandbox-benchmarking

If the underlying physical host has not seen that specific snapshot recently, the infrastructure must pull the heavy state data across the network. This cold restore process completely negates the benefit of snapshotting in the first place.

The Agent Tax

This repository highlights a critical challenge for the next generation of software. AI frameworks like OpenClaw rely on persistent, conversational contexts. If every user message requires waking up a sandbox from a snapshot, a six-second penalty on every interaction breaks the user experience.

The sandbox-benchmarking project proves that infrastructure providers need to rethink their readiness metrics. Until the API accurately reflects when a machine can actually execute code, developers will continue paying this hidden tax.