FutureShow: The AI Battle Arena Where the Future is the Benchmark

How HKUDS is solving the LLM contamination crisis by forcing models to bet against the real-world wisdom of the crowd.

• View on GitHub • More from HKUDS

A black and white illustration of a futuristic coliseum where human silhouettes and server racks face off around a holographic globe.
FutureShow pits frontier models against human prediction markets in real-time.

Key Takeaways

The End of the Open-Book Test

The data contamination crisis has broken AI benchmarking. Frontier models are trained on the entire internet. This means they have likely memorized the answers to static tests like MMLU and GSM8K before the evaluation even begins. It is an open-book test where the student already has the answer key.

FutureShow flips the script by using the only data set an LLM cannot possibly have in its training weights: the future. Built by the HKUDS team, the platform connects directly to Polymarket. It forces models like GPT-5 and DeepSeek to predict the outcome of live geopolitical, economic, and cultural events. If a model wants a high score, it has to accurately model a world that hasn't happened yet.

Inside the Forecasting Loop

To compete against real money, an AI needs more than just its internal weights. FutureShow implements a sophisticated ReAct (Reasoning and Acting) loop within its PolymarketForecastAgent. The agent is not simply asked to guess. It is given a budget of turns and a suite of intelligence-gathering tools.

The agent uses Exa to parse clean text from the web, bypasses paywalls with Google News integrations, and scrapes Twitter and Reddit for real-time market sentiment. To prevent infinite research loops, a StepReminder system monitors the API calls. When the agent nears its limit, the system injects a hard stop, forcing the LLM to synthesize its findings into a binary YES or NO prediction with a stated confidence level.

How a raw Polymarket event is processed into a high-confidence forecast.

Measuring Alpha in a Crowd

Accuracy alone is a poor metric in prediction markets. Guessing that the sun will rise tomorrow yields a 100 percent accuracy rate but zero financial value. FutureShow introduces a 'Prediction Value' metric based on logarithmic scoring. It measures information gain over the consensus.

If a model predicts a 90 percent likelihood for an event that the market already prices at 89 percent, the reward is marginal. However, if a model predicts a 10 percent black swan event that actually occurs, its score skyrockets. This is the search for contrarian alpha. It rewards models that are right exactly when the human crowd is wrong.

A black and white illustration of a single white bird flying against a massive flock of black birds.
FutureShow's scoring system heavily rewards contrarian predictions that defy market consensus.
FeatureTraditional Benchmarks (MMLU)FutureShow
Data FreshnessStatic (High risk of contamination)Live (Zero risk of contamination)
BaselineAverage human test-takerReal-money market consensus
Primary MetricFlat Accuracy (%)Prediction Value (Log-return alpha)
Feedback LoopNoneFinancial market simulation

The Silicon Trading Floor

The most impressive engineering lies in the tool_polymarket_trade.py module. It does not just record a prediction. It simulates the market impact of placing a trade. When an agent buys a position, a Liquidity Overlay applies a decay factor to the simulated order book.

This prevents 'free money' bugs in the simulation. If an AI decides to bet heavily on a low-liquidity market, the simulated price moves against it through slippage. By modeling these financial guardrails, FutureShow ensures that the alpha generated in the simulation could theoretically survive contact with reality.

Human predictions are derived as YES when market probability > 50%, otherwise NO, representing the collective "Wisdom of the Crowd" benchmark

HKUDS/FutureShow Repository, Project Documentation · HKUDS/FutureShow

Static leaderboards are losing their utility as models absorb the internet. By turning the unknown future into a verifiable test set, FutureShow provides a glimpse into the next era of AI evaluation. It is no longer about who has the most parameters. It is about who can see around the corner.