The Agentic Scraper: Inside extract-getnote-articles

How a pragmatic Playwright script and a Claude Code integration bypass the walled garden of GetNote.

6 min read • View on GitHub • More from dontbesilent2025

A vintage mechanical sorting machine separating jagged rocks from uniform gold ingots. It represents the process of converting messy HTML into clean structured Markdown.
The core value proposition: turning noisy platform data into clean, structured knowledge.
Key Takeaways

Building Tools for Agents

Web scrapers are traditionally invoked via CLI flags by human developers. This project shifts the paradigm by including a skill manifest, officially registering itself as a tool for Anthropic's Claude Code. It defines its own parameters, allowing a Large Language Model to dynamically decide when and how to extract a URL based on conversational context.

{
  "name": "extract_getnote",
  "description": "Extract articles from GetNote URLs",
  "parameters": {
    "type": "object",
    "properties": {
      "url": {
        "type": "string",
        "description": "The GetNote URL to extract"
      }
    },
    "required": ["url"]
  }
}

The Borrow a Session Auth Pattern

Logging into modern web platforms programmatically is a nightmare of CAPTCHAs and behavioral tracking. The login.js script bypasses this entirely using a human-in-the-loop strategy. It opens a visible Chromium window, waits for the human to authenticate, detects the dashboard redirect, and saves the persistent session state to a local directory.

The human-agent handoff bypasses bot detection entirely.

Surviving UI Churn with Heuristics

Walled gardens frequently change their CSS classes to break scrapers. Instead of relying on brittle DOM selectors, the extraction script uses a length-based heuristic. By simply collecting text nodes longer than 50 characters, it ignores navigation menus and footers. This ensures it extracts only the core article content.

A realistic human hand holding out a brass skeleton key to a multi-jointed mechanical robotic hand. This illustrates the human-in-the-loop authentication handoff.
The human user handles the messy real-world authentication, handing the keys to the agent.

The Limits of Generalist Extraction

General-purpose extractors like Python's Trafilatura rely on complex tree-pruning algorithms to handle any website. They are powerful but heavy. By hyper-focusing on a single platform and relying on a persistent Playwright context, this tool guarantees higher fidelity for a specific workflow.

Featureextract-getnote-articlesTrafilatura (Python)
Execution ModelAgent-driven SkillDeveloper-driven Library
RuntimeNode.js (Playwright)Python (Requests/LXML)
AuthenticationPersistent Human-in-the-loopHeaders/Cookies injection
Extraction LogicLength heuristics (>50 chars)HTML tree pruning

The Data Sovereignty Mandate

The ultimate goal of the project is converting locked, proprietary knowledge into raw, portable Markdown. It prevents content creators from losing their scripts and research if the platform shuts down, funneling data directly into local tools like Obsidian.

在 Claude Code 中,直接告诉 Claude: "提取这个知识库的文章:https://www.biji.com/subject/xxx/DEFAULT?..."

Install Script, Project Documentation · install.sh