markdownforagents: Turning the Web Into an Action Surface for LLMs
A Cloudflare-native middleware that strips page chrome, extracts clean markdown, and builds a PageMap of real interactables so agents can read and act without drowning in HTML.
- markdownforagents changes the unit of work from a webpage to an action surface by splitting one URL into readable markdown and a semantic PageMap.
- Its real advantage is not cleaner text, but explicit refs and affordances that let agents target interactables without brittle selectors.
- The pipeline is layered, with content negotiation, Cloudflare AI markdown conversion, DOM parsing, and aggressive noise filtering each doing a small job.
- SSRF protection and edge-native deployment are core product features because arbitrary URL fetching is only useful when it is safe to run at scale.
The browser is the wrong abstraction
Web pages are optimized for eyes, not token budgets. If you hand an LLM the raw DOM, you force it to pay for nav bars, cookie banners, footers, and every hidden div that modern sites pile on top of the actual content.
markdownforagents takes a different bet. It treats a URL as something to be decomposed into two machine-friendly surfaces: one for reading, one for acting.
Two outputs, one URL
The first output is markdown with the chrome stripped away. The second is a PageMap, a compact inventory of what the agent can click, type into, or select. That split matters because reasoning and interaction are different jobs.
Convert any URL into agent-ready markdown or structured JSON.
A model that has to infer actions from prose is slower, more brittle, and more expensive than one that can start from explicit affordances. This repo's premise is that the agent should not need to guess where the buttons are hiding.
How the page gets cleaned
The pipeline is not a single model call dressed up as architecture. It starts with content negotiation, trying Accept: text/markdown first. If the target site already speaks markdown, the service can stop early.
When the page still arrives as HTML, Cloudflare's markdown conversion takes over. After that, the repo uses lightweight DOM parsing, metadata extraction, and noise filtering to remove the junk that would waste context window and confuse downstream agents.
curl -H 'Accept: text/markdown' 'https://target.site/article'
curl 'https://worker.example/r?url=https://target.site/article'
curl 'https://worker.example/map?url=https://target.site/article'
The interesting part is the order of operations. It tries the cheapest, most deterministic path first, then escalates only if it needs to. That is exactly what a production tool should do when every extra token and every extra network hop matters.
Why refs matter more than selectors
Most scrapers return links and text. This one assigns refs and affordances to interactables, so an agent can reason in terms of 'click 7' or 'type into 12' instead of constructing brittle CSS selectors.
Most web pages are noisy for AI systems. This project turns a page into: concise markdown content, machine-readable metadata, actionable UI elements (links/buttons/inputs). So agents can read and act with less prompt bloat and fewer parsing hacks.
That is a small design choice with a big payoff. LLMs are better at choosing from a named set of actions than they are at synthesizing DOM paths from scratch.
The security layer is part of the product
Any service that fetches arbitrary URLs has to assume hostile input. markdownforagents blocks private IP ranges, metadata endpoints like 169.254.169.254, and non-web protocols before a request ever reaches the fetch layer.
That is not plumbing. It is the difference between a clever demo and a service you can expose to agents without handing them a network probe. If the product is going to sit between untrusted prompts and the public web, SSRF protection belongs in the headline architecture, not the appendix.
How it stacks up
| Tool | Clean markdown | Action refs | Edge-native | Operational load |
|---|---|---|---|---|
| markdownforagents | Yes | Yes | Yes | Low |
| Cloudflare Markdown for Agents | Yes | No | Yes | Low |
| Readability-style extractors | Yes | No | Sometimes | Low |
| Headless browser + LLM | Sometimes | Sometimes | No | High |
Cloudflare's native markdown conversion is the closest neighbor, but markdownforagents goes further by adding explicit action semantics. Readability-style tools solve extraction. Headless browsers solve interaction. This repo tries to compress both jobs into an edge-friendly middle layer.
Built for agents, not for theater
The repo shape backs up the thesis. The API lives in apps/api, the frontend in apps/web, and the PageMap logic in packages/pagemap, which keeps the transformation engine testable and separate from deployment concerns.
That split suggests a philosophy: keep the core small, keep the edge fast, and make the output legible enough that an agent can act without a browser session doing the thinking for it. alvinunreal is not chasing a prettier reader mode. They are carving a smaller contract between the web and automation.