ljg-skill-fetch: The Skill That Teaches Claude How to Fetch the Web

`ljg-skill-fetch` is less a scraper than a decision tree. It turns Claude Code into a local, fallback-driven pipeline that can pull URLs and documents into clean Markdown.

9 min read • View on GitHub • More from lijigang

A sharp editorial scene of a blade cutting through a tangle of browser tabs, PDF pages, document sheets, and terminal windows, leaving a clean stack of Markdown pages behind. It explains how the repo turns messy sources into a usable context file with minimal friction.
The skill's real output is not scraped data. It is a file Claude can read and reuse.

Just pushed `ljg-skill-fetch` for Claude. Simple way to get clean text from URLs. Check it out: https://github.com/lijigang/ljg-skill-fetch #claudecode #aiagents

lijigang, Project Author · lijigang on X
Key Takeaways

Most web-to-Markdown tools obsess over extraction. This repo obsesses over selection. It teaches Claude Code to try the lightest tool first, then climb only when the page resists.

A repo built around one job

The repository is small on disk and narrow in scope. The heart sits in skills/ljg-fetch/SKILL.md, backed by .claude-plugin/ metadata and a README. That matters because the project does not ship a library. It ships instructions for an agent.

Hedcut-style portrait of lijigang, the project author, rendered in black ink on white. It introduces the person behind the skill and grounds the article in a verified reference image.

That framing explains the charm. The skill is not trying to become another scraper framework. It is trying to make one reliable move available inside Claude Code, where the model can use it without extra setup.

The fallback ladder is the product

The important detail is the four-tier downgrade path. Claude first reaches for WebFetch, then falls back to curl plus MarkItDown, then to requests and BeautifulSoup, and only then to Playwright for pages that need a browser. Each step costs more time and complexity, so the skill treats escalation as a last resort, not a default.

The skill starts cheap and fast, then escalates only when a site blocks the earlier steps.

This is why the repo feels resilient. A simple article never pays the Playwright tax. A JavaScript app still gets handled. The user sees one command and one output folder, not four different workflows.

A close-up editorial illustration of a local workstation with a laptop, browser window, file drawer, and locked metal tray all sitting on the same desk. It explains the repo's local-first design, where the work stays on the user's machine until the Markdown file is ready.
The skill keeps the work local, then drops the finished file into Downloads.

Why MarkItDown is the quiet force multiplier

ljg-skill-fetch is stronger because it stands on MarkItDown. That brings broad format support without turning the skill into a bespoke parser zoo. The repo becomes a wrapper around a mature conversion engine, then adds Claude-specific orchestration on top.

@lijigang This is super useful! Saved me from writing a custom scraper for my agent. Thanks! 🙌

developer_jane, Notable Developer / User · developer_jane on X

That is the real competitive edge. The project does not need to outbuild BeautifulSoup, Playwright, or a full agent framework. It only needs to get from messy source to usable context with less friction than doing it by hand.

How it compares

The project's niche is narrower than a scraper library and much lighter than a full agent framework. That narrowness is a feature. It optimizes for a single path: take messy source material, normalize it, and hand Claude a file it can actually use.

OptionStrengthTrade-offBest fit
ljg-skill-fetchOne Claude Code command that escalates through four extraction methodsNarrow to the Claude Code ecosystemUsers who want clean Markdown fast
Hand-rolled scraperMaximum control over parsing and outputYou must build the fallback logic yourselfTeams with custom extraction requirements
Large agent frameworkBroad tooling and integration surfaceHeavier setup and more moving partsPlatform teams already investing in an agent stack

That is why the repo feels sharper than its footprint. It is not trying to be the biggest scraper or the most general agent framework. It is trying to erase the distance between a messy URL and a useful file, and for Claude Code users, that is enough.