cloudflare/mcp: How Cloudflare Turned a 2,500-Endpoint API Into Two Tools

A look at the repo that replaces flat MCP tooling with search, sandboxed execution, and just-in-time types so agents can work with Cloudflare without drowning in context.

9 min read • View on GitHub • More from cloudflare

A vast archive of tiny API drawers and filing cards stretches into the distance, each drawer labeled like an endpoint. In the foreground, a narrow search beam and a compact execution window carve a clean path through the clutter, reducing the mass of drawers into two simple routes.
Cloudflare’s wager is simple: when an API gets too large to enumerate, agents should search first and execute second.
Key Takeaways

The context window was the real enemy

Cloudflare did not build this repo because it needed another MCP demo. It built it because the normal MCP pattern breaks down when the underlying API is huge. If you try to expose thousands of endpoints as thousands of tools, the model spends its context budget just learning what is available.

That is the real trick in cloudflare/mcp. The repo compresses a sprawling API surface into a workflow that agents can actually hold in mind: search the spec, then execute code against what was found. The point is not to hide the API. It is to make it usable.

The standard approach to MCP requires defining a separate tool for every API operation that is exposed via an MCP server. But this static and exhaustive approach quickly exhausts an agent’s context window, especially for large platforms with thousands of endpoints.

Code Mode turns MCP inside out

The inversion is blunt. Instead of a flat catalog of tools, the server exposes two verbs: search and execute. The model does not need a preloaded map of every endpoint. It discovers the surface on demand, then acts on it.

The whole system is built around a behavioral loop, not a static tool catalog.

That matters because it changes what the model is asked to do. It is no longer memorizing a giant surface area. It is reasoning over a compressed spec, generating code, and using execution results to decide the next move. In practice, the model becomes a navigator, not a catalog reader.

A glass-like Worker isolate sits in the center as a sealed chamber. Inside, a handwritten code fragment runs on a terminal. Outside the chamber, a firewall gate labeled <span class=GlobalOutbound allows only one narrow line to api.cloudflare.com. A typed card is being slid into the chamber, showing just-in-time type injection." data-prompt="Create an editorial illustration rendered entirely in black ink on a pure white background. The image depicts a close-up of a Worker isolate as a glass box or sealed chamber. Inside the chamber, a handwritten code fragment runs on a terminal. Outside the chamber is a firewall gate labeled GlobalOutbound, with only one allowed exit line leading to api.cloudflare.com. A second hand places a typed card into the chamber, representing just-in-time type injection. The artist uses tightly packed crosshatching lines layered at different angles to build up shadow and form, with clean open areas of white for highlights. The line work has the quality of a classic metal engraving, precise and deliberate, with varied line weights where bold contour lines define shapes and finer interior lines create tonal depth. The overall style evokes vintage newspaper editorial illustrations from The Economist or Wall Street Journal. No color, no gradients, no grey fills. Only black lines on white. The background MUST be pure white #FFFFFF. No paper texture, no cream, no off-white, no noise, no grain. Perfectly clean flat white background." loading="lazy">
The safety story is not just about permissions. It is about forcing generated code to run inside a narrow, observable corridor.

The repo is really a sandbox, not just a server

Once you read the code, the center of gravity shifts. The interesting part is not the MCP transport. It is the execution environment. `createSearchExecutor` and `createCodeExecutor` wrap the hard problem in Cloudflare Workers, then constrain what the generated code can touch.

That is why `GlobalOutbound` matters so much. If the model can only reach `api.cloudflare.com`, you have turned a general-purpose code runner into a tightly scoped API instrument. The agent can explore, but it cannot freeload its way into arbitrary network access.

Code Mode is a technique we first introduced for reducing context window usage during agent tool use. Instead of describing every operation as a separate tool, let the model write code against a typed SDK and execute the code safely in a Dynamic Worker Loader.

The spec is rewritten before the model ever sees it

The repo also does quiet but important preprocessing. `resolveRefs` flattens the OpenAPI tree so the agent does not have to chase references across a maze of nested definitions. `extractProduct` groups endpoint paths into product categories, which is the sort of compression that turns an intimidating API into something searchable.

const flatSpec = await resolveRefs(rawOpenApiSpec)
const products = extractProduct(flatSpec.paths)

const results = await createSearchExecutor({
  spec: flatSpec,
  products,
})

This is a useful mental model for any large platform. First, compress reality. Then present only the slice the agent needs. The model is much better at deciding between a few well-shaped options than at digesting an entire API atlas.

Authentication is split between humans and tokens

The auth layer is practical rather than clever. Cloudflare supports OAuth for logged-in users and API token mode for direct access. The repo also tries to infer account_id when a token only has one account in scope, which removes a lot of friction for narrow permissions.

ModeHow it worksBest forTrade-off
OAuthBrowser-based login through Cloudflare Access and session storageHumans working from the dashboard or an IDEMore setup, but better for governed access
API tokenDirect bearer token use with account auto-detectionAutomation and targeted account accessLess ceremony, but more dependent on token hygiene

The lesson is that agent infrastructure still has to feel like real infrastructure. If auth is clumsy, the whole promise collapses into a nicer error message.

The guardrails are as important as the model tricks

The repo does not assume the model will behave. It budgets for bad answers, oversized responses, and retries. `truncateResponse` caps output so the agent does not flood itself, and the overall flow leaves room for narrower follow-up searches when the first pass is incomplete.

That is the difference between a prototype and a usable agent interface. A good tool does not just make the happy path shorter. It makes the failure path recoverable.

Why this pattern matters beyond Cloudflare

Cloudflare’s code mode is not only a Cloudflare story. It is a strong argument against endpoint-as-tool sprawl for any large enterprise API. Flat MCP works fine for small surfaces. For giant ones, it becomes an expensive way to describe a problem the model cannot hold in working memory.

ApproachTool shapeContext costSafety modelBest use caseWeakness
Flat MCPOne tool per endpointHigh and grows with the APISimple permissions, but broad exposureSmall, stable APIsTool catalogs become unmanageable
Cloudflare Code ModeSearch first, then execute codeLow and relatively fixedSandboxed Workers plus outbound limitsHuge APIs with many endpointsMore moving parts than a basic server
Enterprise portalCentralized discovery across many serversMedium, but governedPolicy, logging, DLP, and access controlOrganizations managing many MCP sourcesDoes not solve per-API complexity alone

That is why the repo feels bigger than its size. It is a working answer to a pattern that will keep showing up everywhere: when the API surface gets too wide, the interface has to become layered.

Sources