FutureShow: The AI Battle Arena Where the Future is the Benchmark
How HKUDS is solving the LLM contamination crisis by forcing models to bet against the real-world wisdom of the crowd.
- FutureShow eliminates data contamination by testing models on live events from real-world prediction markets.
- The platform uses a ReAct loop to equip agents with real-time intelligence tools like Exa and social media scrapers.
- A logarithmic scoring system rewards models for finding contrarian alpha against the human crowd consensus.
- The simulation includes liquidity decay factors to model the actual financial impact of placing trades in thin markets.
The End of the Open-Book Test
The data contamination crisis has broken AI benchmarking. Frontier models are trained on the entire internet. This means they have likely memorized the answers to static tests like MMLU and GSM8K before the evaluation even begins. It is an open-book test where the student already has the answer key.
FutureShow flips the script by using the only data set an LLM cannot possibly have in its training weights: the future. Built by the HKUDS team, the platform connects directly to Polymarket. It forces models like GPT-5 and DeepSeek to predict the outcome of live geopolitical, economic, and cultural events. If a model wants a high score, it has to accurately model a world that hasn't happened yet.
Inside the Forecasting Loop
To compete against real money, an AI needs more than just its internal weights. FutureShow implements a sophisticated ReAct (Reasoning and Acting) loop within its PolymarketForecastAgent. The agent is not simply asked to guess. It is given a budget of turns and a suite of intelligence-gathering tools.
The agent uses Exa to parse clean text from the web, bypasses paywalls with Google News integrations, and scrapes Twitter and Reddit for real-time market sentiment. To prevent infinite research loops, a StepReminder system monitors the API calls. When the agent nears its limit, the system injects a hard stop, forcing the LLM to synthesize its findings into a binary YES or NO prediction with a stated confidence level.
Measuring Alpha in a Crowd
Accuracy alone is a poor metric in prediction markets. Guessing that the sun will rise tomorrow yields a 100 percent accuracy rate but zero financial value. FutureShow introduces a 'Prediction Value' metric based on logarithmic scoring. It measures information gain over the consensus.
If a model predicts a 90 percent likelihood for an event that the market already prices at 89 percent, the reward is marginal. However, if a model predicts a 10 percent black swan event that actually occurs, its score skyrockets. This is the search for contrarian alpha. It rewards models that are right exactly when the human crowd is wrong.
| Feature | Traditional Benchmarks (MMLU) | FutureShow |
|---|---|---|
| Data Freshness | Static (High risk of contamination) | Live (Zero risk of contamination) |
| Baseline | Average human test-taker | Real-money market consensus |
| Primary Metric | Flat Accuracy (%) | Prediction Value (Log-return alpha) |
| Feedback Loop | None | Financial market simulation |
The Silicon Trading Floor
The most impressive engineering lies in the tool_polymarket_trade.py module. It does not just record a prediction. It simulates the market impact of placing a trade. When an agent buys a position, a Liquidity Overlay applies a decay factor to the simulated order book.
This prevents 'free money' bugs in the simulation. If an AI decides to bet heavily on a low-liquidity market, the simulated price moves against it through slippage. By modeling these financial guardrails, FutureShow ensures that the alpha generated in the simulation could theoretically survive contact with reality.
Human predictions are derived as YES when market probability > 50%, otherwise NO, representing the collective "Wisdom of the Crowd" benchmark
Static leaderboards are losing their utility as models absorb the internet. By turning the unknown future into a verifiable test set, FutureShow provides a glimpse into the next era of AI evaluation. It is no longer about who has the most parameters. It is about who can see around the corner.