MindSpider: The Crawler That Asks an LLM What to Search First
A search-then-crawl system that turns 13 hot lists into targeted scraping jobs, then digs deep into posts, comments, and sentiment across Chinese platforms.
- MindSpider treats topic discovery as the first automation problem, so the crawler spends its effort on the right questions.
- Its LLM stage compresses noisy hot lists into normalized search terms and summaries that can drive deep crawling.
- The code is built for production messiness, with fallback parsing, browser automation, and daily overwrite refreshes.
- Against MediaCrawler, Scrapy, and Crawlee, MindSpider's real edge is upstream intelligence, not just scraping mechanics.
The crawler that chooses its own keywords
Most scrapers start with a URL list or a keyword list. MindSpider starts with a judgment call. It collects 13 hot lists, asks an LLM to sort the signal from the noise, and only then decides what deserves a deep crawl.
That move changes the job. The system is not just fetching content faster. It is deciding what to search for in the first place, which is the part most public opinion workflows still leave to a human.
MindSpider:专为舆情分析设计的AI爬虫
Why manual keyword scraping breaks down
Public attention fragments quickly. A topic on Weibo may surface with one phrase, on Zhihu with another, and on GitHub Trending with a third. If you hand-author the keyword list, you are already behind the curve.
MindSpider is built for that gap. It treats the hot list itself as raw material, then uses DeepSeek to synthesize search terms that are broad enough to catch variation but narrow enough to stay useful.
That matters because the system is not chasing archival completeness. It is trying to keep a daily pulse on what people are actually talking about, which means freshness matters more than elegance.
Search first, crawl second
This is the core loop: noisy hot lists go in, normalized topics come out, and those topics become crawl jobs. The discovery layer is not a preface. It is the engine that makes the rest of the pipeline worth running.
In the repository, that idea is split across two modules. BroadTopicExtraction gathers the hot lists and asks the model to produce structured topics. DeepSentimentCrawling consumes those topics and pushes them into a modified MediaCrawler stack for deeper collection.
Under the hood
get_today_news.py pulls hot lists from a shared API and normalizes 13 sources into one schema. topic_extractor.py then builds a prompt for DeepSeek and expects a JSON object back, not a chatty answer. If the model returns messy output, a manual parser tries to recover the keywords instead of failing the run.
raw_hot_lists = get_today_news()
analysis = topic_extractor.extract(raw_hot_lists)
keywords = analysis.get('keywords', [])
for keyword in keywords:
crawl_result = platform_crawler.run(keyword)
database_manager.overwrite_daily_topics(crawl_result)
That fallback logic is the tell. MindSpider is using an LLM, but it does not trust the LLM blindly. It treats model output as structured input that may need repair, which is the right instinct for a production crawler.
The crawl side is equally practical. The system uses Playwright for browser automation, AsyncIO for concurrency, and a database overwrite pattern for daily refreshes, so the latest topic snapshot replaces the old one instead of piling up duplicates.
MindSpider vs the usual scraping stack
| Project | Starting point | Discovery layer | Crawl style | Best at |
|---|---|---|---|---|
| MindSpider | 13 hot lists and AI-synthesized topics | Yes, LLM-driven topic extraction | Browser automation plus platform jobs | Public opinion workflows |
| MediaCrawler | Manual keyword input | No | Browser scraping | Deep extraction on supported platforms |
| Scrapy | URLs and parsing rules | No | HTTP-first spiders | Custom crawls at scale |
| Crawlee | URLs or search terms | No | HTTP and browser hybrid | General-purpose automation |
MediaCrawler is the engine MindSpider inherits. The difference is that MindSpider adds a decision layer above it, so the crawler is not waiting for a human to name the topic.
Scrapy and Crawlee are strong general tools, but they leave topic selection to the operator. MindSpider narrows the use case to public opinion analysis and uses that constraint to automate the most expensive part of the workflow.
What this pattern suggests
The bigger shift is upstream. Crawlers used to be judged by how well they fetched content. Systems like MindSpider are starting to be judged by how well they choose their inputs.
That is a meaningful change for research teams, PR analysts, and anyone tracking fast-moving conversation. If the signal is hidden in the topic selection step, then the crawler that can think about search before it scrapes is the one that gets there first.