last30days-skill: The High-Stakes Architecture of Real-Time Research

How a modular signal engine triangulates truth across prediction markets, social sentiment, and the open web to kill the LLM knowledge cutoff.

• View on GitHub • More from mvanhorn

A heavy anchor labeled Polymarket holding down chaotic balloons labeled X, Reddit, and TikTok, preventing them from floating into a storm cloud of hype.
Financial stakes ground social sentiment, preventing hype from overwhelming the research engine.

Key Takeaways

The Financialization of Fact

The knowledge cutoff is the original sin of Large Language Models. While most developers attempt to fix this with Retrieval-Augmented Generation (RAG) or basic Google Search tools, last30days-skill treats the internet not as a library, but as a high-frequency trading floor.

By treating Polymarket as a primary signal alongside Reddit and YouTube, this project codifies the search habits of an investigative researcher into a deterministic Python engine. It is a deep research tool that values skin in the game over SEO-optimized text.

AI agent skill that researches any topic across Reddit, X, YouTube, HN, Polymarket, and the web - then synthesizes a grounded summary

mvanhorn, Project Creator and Lead Maintainer · GitHub - mvanhorn/last30days-skill

The Intent-Aware Router

Instead of hitting every API for every query, the system uses a heuristic intent classifier. The query_type.py module programmatically decides that a product search requires YouTube transcripts and Reddit threads, while a concept search relies on blog posts and Hacker News.

How the Intent Router selectively activates data sources based on query classification.

Scoring the Noise

Raw API results are invariably noisy. The score.py module implements a multi-factor scoring algorithm to rank items by relevance, recency, and engagement. Crucially, it uses logarithmic scaling for engagement metrics.

This mathematical approach ensures a viral tweet with one million likes doesn't completely drown out a highly relevant Reddit comment with one hundred upvotes. Truth, in this engine, is a weighted average of convergence across platforms.

Featurelast30days-skillStandard RAG
Primary SignalSocial/Financial (Reddit, Polymarket)SEO/Web (Google, Bing)
Temporal WindowHard 30-day decay curveGeneral or arbitrary
Truth MechanismCross-platform triangulationLLM summarization of top links

Beyond the Web: The Social Scrapers

To access unfiltered discourse, the project bypasses official, highly restricted APIs. It uses tools like yt-dlp for YouTube transcripts and ScrapeCreators for Reddit and TikTok, extracting raw human sentiment before it is packaged by search engine algorithms.

A mechanical sieve catching identical gears while allowing a single unique master gear to pass through.
The deduplication engine filters out echo-chamber repetitions of the same viral story.

Finally, a hybrid deduplication system uses token Jaccard similarity to prevent the echo chamber effect, ensuring the final synthesized report provides a diverse, high-signal briefing of the modern internet's true pulse.