markdownforagents: Turning the Web Into an Action Surface for LLMs

A Cloudflare-native middleware that strips page chrome, extracts clean markdown, and builds a PageMap of real interactables so agents can read and act without drowning in HTML.

8 min read • View on GitHub • More from alvinunreal

A noisy browser page is pressed through a machine and split into two clean outputs, one reading layer and one action map. The image explains the core idea that a single web page can become both readable markdown and a semantic map of things an agent can do.
One URL, two outputs: a clean reading surface and a compact action surface.
Key Takeaways

The browser is the wrong abstraction

Web pages are optimized for eyes, not token budgets. If you hand an LLM the raw DOM, you force it to pay for nav bars, cookie banners, footers, and every hidden div that modern sites pile on top of the actual content.

markdownforagents takes a different bet. It treats a URL as something to be decomposed into two machine-friendly surfaces: one for reading, one for acting.

Two outputs, one URL

The first output is markdown with the chrome stripped away. The second is a PageMap, a compact inventory of what the agent can click, type into, or select. That split matters because reasoning and interaction are different jobs.

Convert any URL into agent-ready markdown or structured JSON.

alvinunreal, Project Creator and Maintainer · markdownforagents README

A model that has to infer actions from prose is slower, more brittle, and more expensive than one that can start from explicit affordances. This repo's premise is that the agent should not need to guess where the buttons are hiding.

A single URL becomes both readable markdown and a semantic map of actions.

How the page gets cleaned

A tight view of a web page where cookie banners, nav bars, and footer clutter are being pulled away like debris through a sieve. In the center, the actual content remains, with clean numbered refs attached to the interactive elements to show what survives the filtering step.
Noise goes first. The useful content and interactables stay behind.

The pipeline is not a single model call dressed up as architecture. It starts with content negotiation, trying Accept: text/markdown first. If the target site already speaks markdown, the service can stop early.

When the page still arrives as HTML, Cloudflare's markdown conversion takes over. After that, the repo uses lightweight DOM parsing, metadata extraction, and noise filtering to remove the junk that would waste context window and confuse downstream agents.

curl -H 'Accept: text/markdown' 'https://target.site/article'
curl 'https://worker.example/r?url=https://target.site/article'
curl 'https://worker.example/map?url=https://target.site/article'

The interesting part is the order of operations. It tries the cheapest, most deterministic path first, then escalates only if it needs to. That is exactly what a production tool should do when every extra token and every extra network hop matters.

Why refs matter more than selectors

A drafting table shows a page blueprint with several interactive nodes pinned in place, each carrying a small ref number and an affordance tag such as click, type, or select. The image explains that the project assigns stable action labels instead of asking an agent to invent CSS selectors.
Refs give the agent a stable target. Affordances tell it what kind of action belongs there.

Most scrapers return links and text. This one assigns refs and affordances to interactables, so an agent can reason in terms of 'click 7' or 'type into 12' instead of constructing brittle CSS selectors.

Most web pages are noisy for AI systems. This project turns a page into: concise markdown content, machine-readable metadata, actionable UI elements (links/buttons/inputs). So agents can read and act with less prompt bloat and fewer parsing hacks.

alvinunreal, Project Creator and Maintainer · markdownforagents README

That is a small design choice with a big payoff. LLMs are better at choosing from a named set of actions than they are at synthesizing DOM paths from scratch.

The security layer is part of the product

Any service that fetches arbitrary URLs has to assume hostile input. markdownforagents blocks private IP ranges, metadata endpoints like 169.254.169.254, and non-web protocols before a request ever reaches the fetch layer.

That is not plumbing. It is the difference between a clever demo and a service you can expose to agents without handing them a network probe. If the product is going to sit between untrusted prompts and the public web, SSRF protection belongs in the headline architecture, not the appendix.

How it stacks up

A split scene contrasts a heavy, cluttered browser stack on the left with a light edge pipeline on the right. The left side feels layered and expensive, while the right side turns a single URL into markdown and a PageMap with far less machinery, showing the tradeoff between traditional scraping and the repo's approach.
The alternative is still the old stack: a browser, a parser, and a lot of orchestration.
ToolClean markdownAction refsEdge-nativeOperational load
markdownforagentsYesYesYesLow
Cloudflare Markdown for AgentsYesNoYesLow
Readability-style extractorsYesNoSometimesLow
Headless browser + LLMSometimesSometimesNoHigh

Cloudflare's native markdown conversion is the closest neighbor, but markdownforagents goes further by adding explicit action semantics. Readability-style tools solve extraction. Headless browsers solve interaction. This repo tries to compress both jobs into an edge-friendly middle layer.

Built for agents, not for theater

The repo shape backs up the thesis. The API lives in apps/api, the frontend in apps/web, and the PageMap logic in packages/pagemap, which keeps the transformation engine testable and separate from deployment concerns.

That split suggests a philosophy: keep the core small, keep the edge fast, and make the output legible enough that an agent can act without a browser session doing the thinking for it. alvinunreal is not chasing a prettier reader mode. They are carving a smaller contract between the web and automation.