MindSpider: The Crawler That Asks an LLM What to Search First

A search-then-crawl system that turns 13 hot lists into targeted scraping jobs, then digs deep into posts, comments, and sentiment across Chinese platforms.

9 min read • View on GitHub • More from 666ghj

A wide field of hot-list placards feeds into a central mechanical sieve, which outputs a neat stack of keyword cards and a waiting crawler cart. It explains how MindSpider converts trend noise into a search plan before it starts scraping.
MindSpider does not begin with a target list. It begins with noisy signals, then turns them into crawlable intent.
Key Takeaways

The crawler that chooses its own keywords

Most scrapers start with a URL list or a keyword list. MindSpider starts with a judgment call. It collects 13 hot lists, asks an LLM to sort the signal from the noise, and only then decides what deserves a deep crawl.

That move changes the job. The system is not just fetching content faster. It is deciding what to search for in the first place, which is the part most public opinion workflows still leave to a human.

MindSpider:专为舆情分析设计的AI爬虫

666ghj, Author/Maintainer · MindSpider README

Why manual keyword scraping breaks down

Public attention fragments quickly. A topic on Weibo may surface with one phrase, on Zhihu with another, and on GitHub Trending with a third. If you hand-author the keyword list, you are already behind the curve.

MindSpider is built for that gap. It treats the hot list itself as raw material, then uses DeepSeek to synthesize search terms that are broad enough to catch variation but narrow enough to stay useful.

That matters because the system is not chasing archival completeness. It is trying to keep a daily pulse on what people are actually talking about, which means freshness matters more than elegance.

Search first, crawl second

This is the core loop: noisy hot lists go in, normalized topics come out, and those topics become crawl jobs. The discovery layer is not a preface. It is the engine that makes the rest of the pipeline worth running.

The important transition is upstream. MindSpider turns trend discovery into crawl inputs, then routes those inputs into platform-specific extraction and daily storage.

In the repository, that idea is split across two modules. BroadTopicExtraction gathers the hot lists and asks the model to produce structured topics. DeepSentimentCrawling consumes those topics and pushes them into a modified MediaCrawler stack for deeper collection.

Under the hood

get_today_news.py pulls hot lists from a shared API and normalizes 13 sources into one schema. topic_extractor.py then builds a prompt for DeepSeek and expects a JSON object back, not a chatty answer. If the model returns messy output, a manual parser tries to recover the keywords instead of failing the run.

raw_hot_lists = get_today_news()
analysis = topic_extractor.extract(raw_hot_lists)
keywords = analysis.get('keywords', [])

for keyword in keywords:
    crawl_result = platform_crawler.run(keyword)
    database_manager.overwrite_daily_topics(crawl_result)

That fallback logic is the tell. MindSpider is using an LLM, but it does not trust the LLM blindly. It treats model output as structured input that may need repair, which is the right instinct for a production crawler.

The crawl side is equally practical. The system uses Playwright for browser automation, AsyncIO for concurrency, and a database overwrite pattern for daily refreshes, so the latest topic snapshot replaces the old one instead of piling up duplicates.

MindSpider vs the usual scraping stack

ProjectStarting pointDiscovery layerCrawl styleBest at
MindSpider13 hot lists and AI-synthesized topicsYes, LLM-driven topic extractionBrowser automation plus platform jobsPublic opinion workflows
MediaCrawlerManual keyword inputNoBrowser scrapingDeep extraction on supported platforms
ScrapyURLs and parsing rulesNoHTTP-first spidersCustom crawls at scale
CrawleeURLs or search termsNoHTTP and browser hybridGeneral-purpose automation

MediaCrawler is the engine MindSpider inherits. The difference is that MindSpider adds a decision layer above it, so the crawler is not waiting for a human to name the topic.

Scrapy and Crawlee are strong general tools, but they leave topic selection to the operator. MindSpider narrows the use case to public opinion analysis and uses that constraint to automate the most expensive part of the workflow.

What this pattern suggests

The bigger shift is upstream. Crawlers used to be judged by how well they fetched content. Systems like MindSpider are starting to be judged by how well they choose their inputs.

That is a meaningful change for research teams, PR analysts, and anyone tracking fast-moving conversation. If the signal is hidden in the topic selection step, then the crawler that can think about search before it scrapes is the one that gets there first.