The Agentic Scraper: Inside extract-getnote-articles
How a pragmatic Playwright script and a Claude Code integration bypass the walled garden of GetNote.
- The project operates as a Claude Code Skill, shifting execution from human CLI commands to LLM orchestration.
- It bypasses complex bot detection by using a persistent browser session initially authenticated by a human user.
- Length-based text heuristics replace brittle CSS selectors to ensure extraction survives platform UI updates.
- The ultimate goal is data sovereignty, converting locked proprietary content into portable local Markdown files.
Building Tools for Agents
Web scrapers are traditionally invoked via CLI flags by human developers. This project shifts the paradigm by including a skill manifest, officially registering itself as a tool for Anthropic's Claude Code. It defines its own parameters, allowing a Large Language Model to dynamically decide when and how to extract a URL based on conversational context.
{
"name": "extract_getnote",
"description": "Extract articles from GetNote URLs",
"parameters": {
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "The GetNote URL to extract"
}
},
"required": ["url"]
}
}
The Borrow a Session Auth Pattern
Logging into modern web platforms programmatically is a nightmare of CAPTCHAs and behavioral tracking. The login.js script bypasses this entirely using a human-in-the-loop strategy. It opens a visible Chromium window, waits for the human to authenticate, detects the dashboard redirect, and saves the persistent session state to a local directory.
Surviving UI Churn with Heuristics
Walled gardens frequently change their CSS classes to break scrapers. Instead of relying on brittle DOM selectors, the extraction script uses a length-based heuristic. By simply collecting text nodes longer than 50 characters, it ignores navigation menus and footers. This ensures it extracts only the core article content.
The Limits of Generalist Extraction
General-purpose extractors like Python's Trafilatura rely on complex tree-pruning algorithms to handle any website. They are powerful but heavy. By hyper-focusing on a single platform and relying on a persistent Playwright context, this tool guarantees higher fidelity for a specific workflow.
| Feature | extract-getnote-articles | Trafilatura (Python) |
|---|---|---|
| Execution Model | Agent-driven Skill | Developer-driven Library |
| Runtime | Node.js (Playwright) | Python (Requests/LXML) |
| Authentication | Persistent Human-in-the-loop | Headers/Cookies injection |
| Extraction Logic | Length heuristics (>50 chars) | HTML tree pruning |
The Data Sovereignty Mandate
The ultimate goal of the project is converting locked, proprietary knowledge into raw, portable Markdown. It prevents content creators from losing their scripts and research if the platform shuts down, funneling data directly into local tools like Obsidian.
在 Claude Code 中,直接告诉 Claude: "提取这个知识库的文章:https://www.biji.com/subject/xxx/DEFAULT?..."