cloudflare/mcp: How Cloudflare Turned a 2,500-Endpoint API Into Two Tools
A look at the repo that replaces flat MCP tooling with search, sandboxed execution, and just-in-time types so agents can work with Cloudflare without drowning in context.
- `cloudflare/mcp` argues that giant APIs should not be flattened into giant tool lists for agents.
- Its core move is to split interaction into discovery and execution, which keeps context small and behavior tractable.
- The repo is as much a sandbox as a server, with outbound controls and dynamic code loading doing the safety work.
- The bigger lesson is that MCP for enterprise-scale systems may need layered, code-first interfaces instead of endpoint sprawl.
The context window was the real enemy
Cloudflare did not build this repo because it needed another MCP demo. It built it because the normal MCP pattern breaks down when the underlying API is huge. If you try to expose thousands of endpoints as thousands of tools, the model spends its context budget just learning what is available.
That is the real trick in cloudflare/mcp. The repo compresses a sprawling API surface into a workflow that agents can actually hold in mind: search the spec, then execute code against what was found. The point is not to hide the API. It is to make it usable.
The standard approach to MCP requires defining a separate tool for every API operation that is exposed via an MCP server. But this static and exhaustive approach quickly exhausts an agent’s context window, especially for large platforms with thousands of endpoints.
Code Mode turns MCP inside out
The inversion is blunt. Instead of a flat catalog of tools, the server exposes two verbs: search and execute. The model does not need a preloaded map of every endpoint. It discovers the surface on demand, then acts on it.
That matters because it changes what the model is asked to do. It is no longer memorizing a giant surface area. It is reasoning over a compressed spec, generating code, and using execution results to decide the next move. In practice, the model becomes a navigator, not a catalog reader.
GlobalOutbound allows only one narrow line to api.cloudflare.com. A typed card is being slid into the chamber, showing just-in-time type injection." data-prompt="Create an editorial illustration rendered entirely in black ink on a pure white background. The image depicts a close-up of a Worker isolate as a glass box or sealed chamber. Inside the chamber, a handwritten code fragment runs on a terminal. Outside the chamber is a firewall gate labeled GlobalOutbound, with only one allowed exit line leading to api.cloudflare.com. A second hand places a typed card into the chamber, representing just-in-time type injection. The artist uses tightly packed crosshatching lines layered at different angles to build up shadow and form, with clean open areas of white for highlights. The line work has the quality of a classic metal engraving, precise and deliberate, with varied line weights where bold contour lines define shapes and finer interior lines create tonal depth. The overall style evokes vintage newspaper editorial illustrations from The Economist or Wall Street Journal. No color, no gradients, no grey fills. Only black lines on white. The background MUST be pure white #FFFFFF. No paper texture, no cream, no off-white, no noise, no grain. Perfectly clean flat white background." loading="lazy">
The repo is really a sandbox, not just a server
Once you read the code, the center of gravity shifts. The interesting part is not the MCP transport. It is the execution environment. `createSearchExecutor` and `createCodeExecutor` wrap the hard problem in Cloudflare Workers, then constrain what the generated code can touch.
That is why `GlobalOutbound` matters so much. If the model can only reach `api.cloudflare.com`, you have turned a general-purpose code runner into a tightly scoped API instrument. The agent can explore, but it cannot freeload its way into arbitrary network access.
Code Mode is a technique we first introduced for reducing context window usage during agent tool use. Instead of describing every operation as a separate tool, let the model write code against a typed SDK and execute the code safely in a Dynamic Worker Loader.
The spec is rewritten before the model ever sees it
The repo also does quiet but important preprocessing. `resolveRefs` flattens the OpenAPI tree so the agent does not have to chase references across a maze of nested definitions. `extractProduct` groups endpoint paths into product categories, which is the sort of compression that turns an intimidating API into something searchable.
const flatSpec = await resolveRefs(rawOpenApiSpec)
const products = extractProduct(flatSpec.paths)
const results = await createSearchExecutor({
spec: flatSpec,
products,
})
This is a useful mental model for any large platform. First, compress reality. Then present only the slice the agent needs. The model is much better at deciding between a few well-shaped options than at digesting an entire API atlas.
Authentication is split between humans and tokens
The auth layer is practical rather than clever. Cloudflare supports OAuth for logged-in users and API token mode for direct access. The repo also tries to infer account_id when a token only has one account in scope, which removes a lot of friction for narrow permissions.
| Mode | How it works | Best for | Trade-off |
|---|---|---|---|
| OAuth | Browser-based login through Cloudflare Access and session storage | Humans working from the dashboard or an IDE | More setup, but better for governed access |
| API token | Direct bearer token use with account auto-detection | Automation and targeted account access | Less ceremony, but more dependent on token hygiene |
The lesson is that agent infrastructure still has to feel like real infrastructure. If auth is clumsy, the whole promise collapses into a nicer error message.
The guardrails are as important as the model tricks
The repo does not assume the model will behave. It budgets for bad answers, oversized responses, and retries. `truncateResponse` caps output so the agent does not flood itself, and the overall flow leaves room for narrower follow-up searches when the first pass is incomplete.
That is the difference between a prototype and a usable agent interface. A good tool does not just make the happy path shorter. It makes the failure path recoverable.
Why this pattern matters beyond Cloudflare
Cloudflare’s code mode is not only a Cloudflare story. It is a strong argument against endpoint-as-tool sprawl for any large enterprise API. Flat MCP works fine for small surfaces. For giant ones, it becomes an expensive way to describe a problem the model cannot hold in working memory.
| Approach | Tool shape | Context cost | Safety model | Best use case | Weakness |
|---|---|---|---|---|---|
| Flat MCP | One tool per endpoint | High and grows with the API | Simple permissions, but broad exposure | Small, stable APIs | Tool catalogs become unmanageable |
| Cloudflare Code Mode | Search first, then execute code | Low and relatively fixed | Sandboxed Workers plus outbound limits | Huge APIs with many endpoints | More moving parts than a basic server |
| Enterprise portal | Centralized discovery across many servers | Medium, but governed | Policy, logging, DLP, and access control | Organizations managing many MCP sources | Does not solve per-API complexity alone |
That is why the repo feels bigger than its size. It is a working answer to a pattern that will keep showing up everywhere: when the API surface gets too wide, the interface has to become layered.