ljg-skill-fetch: The Skill That Teaches Claude How to Fetch the Web
`ljg-skill-fetch` is less a scraper than a decision tree. It turns Claude Code into a local, fallback-driven pipeline that can pull URLs and documents into clean Markdown.

Just pushed `ljg-skill-fetch` for Claude. Simple way to get clean text from URLs. Check it out: https://github.com/lijigang/ljg-skill-fetch #claudecode #aiagents
- `ljg-skill-fetch` is a manual for Claude Code that turns source gathering into an agent-native workflow.
- Its fallback ladder is the product, because it chooses the lightest extraction path that still works.
- MarkItDown plus browser automation gives the skill breadth without forcing the user into a framework.
- The project wins by shrinking context acquisition to one local command and one clean Markdown file.
Most web-to-Markdown tools obsess over extraction. This repo obsesses over selection. It teaches Claude Code to try the lightest tool first, then climb only when the page resists.
A repo built around one job
The repository is small on disk and narrow in scope. The heart sits in skills/ljg-fetch/SKILL.md, backed by .claude-plugin/ metadata and a README. That matters because the project does not ship a library. It ships instructions for an agent.
That framing explains the charm. The skill is not trying to become another scraper framework. It is trying to make one reliable move available inside Claude Code, where the model can use it without extra setup.
The fallback ladder is the product
The important detail is the four-tier downgrade path. Claude first reaches for WebFetch, then falls back to curl plus MarkItDown, then to requests and BeautifulSoup, and only then to Playwright for pages that need a browser. Each step costs more time and complexity, so the skill treats escalation as a last resort, not a default.
This is why the repo feels resilient. A simple article never pays the Playwright tax. A JavaScript app still gets handled. The user sees one command and one output folder, not four different workflows.
Why MarkItDown is the quiet force multiplier
ljg-skill-fetch is stronger because it stands on MarkItDown. That brings broad format support without turning the skill into a bespoke parser zoo. The repo becomes a wrapper around a mature conversion engine, then adds Claude-specific orchestration on top.
@lijigang This is super useful! Saved me from writing a custom scraper for my agent. Thanks! 🙌
That is the real competitive edge. The project does not need to outbuild BeautifulSoup, Playwright, or a full agent framework. It only needs to get from messy source to usable context with less friction than doing it by hand.
How it compares
The project's niche is narrower than a scraper library and much lighter than a full agent framework. That narrowness is a feature. It optimizes for a single path: take messy source material, normalize it, and hand Claude a file it can actually use.
| Option | Strength | Trade-off | Best fit |
|---|---|---|---|
| ljg-skill-fetch | One Claude Code command that escalates through four extraction methods | Narrow to the Claude Code ecosystem | Users who want clean Markdown fast |
| Hand-rolled scraper | Maximum control over parsing and output | You must build the fallback logic yourself | Teams with custom extraction requirements |
| Large agent framework | Broad tooling and integration surface | Heavier setup and more moving parts | Platform teams already investing in an agent stack |
That is why the repo feels sharper than its footprint. It is not trying to be the biggest scraper or the most general agent framework. It is trying to erase the distance between a messy URL and a useful file, and for Claude Code users, that is enough.