market-ai-resolution: The AI Oracle That Knows When Not to Decide
A TypeScript proof of concept for prediction-market settlement that pulls live evidence, follows market rules, and falls back to UNKNOWN when confidence is too low.
- The repo's core move is to treat uncertainty as a valid settlement outcome, not a bug to hide.
- Live search, structured tools, and source capture turn market rules into a bounded workflow instead of a free-form opinion.
- The confidence threshold matters more than the prompt because it decides when the system must refuse to guess.
- Compared with token-weighted oracle politics, the design is faster and cheaper, but still centralized and experimental.
The machine that can say UNKNOWN
The surprising thing about market-ai-resolution is not that it uses an LLM. It is that the system is built to stop the model from bluffing. When confidence is too low, the resolver does not improvise. It falls back to UNKNOWN and leaves the market unsettled rather than pretending certainty is there.
A proof-of-concept for replacing Polymarket's capital-weighted oracle voting with neutral, evidence-based AI resolution.
That sounds simple until you think about what settlement actually means. A market close is a claim about reality, but reality arrives through messy news reports, partial evidence, and rules that can be more specific than common sense. If the oracle layer is sloppy, the whole market inherits that sloppiness.
Why prediction-market resolution is such a messy job
Prediction markets work when the settlement layer is trusted, fast, and hard to game. Traditional dispute systems often lean on capital-weighted incentives, which makes the final answer expensive to challenge and politically easier to dominate than a casual user would expect. This repo attacks that friction by moving the decision into a rule-bound machine that can gather live evidence and cite what it used.
That is a different claim from predicting markets. Prediction is about forecasting. Resolution is about judgment under a fixed rulebook. The repo is interesting because it treats judgment as an engineering problem.
Inside the resolution loop
The codebase is organized like a small service, not a prompt demo. src/scheduler watches for closed markets, src/resolution prevents duplicate work with a processing lock, src/ai builds the prompt and tool schema, src/providers abstracts market data, and src/events broadcasts completed resolutions through a bus. The public surface sits behind Hono, while validation lives in Zod and the runtime runs on Bun.
type ResolutionResult = {
outcome: string
confidence: number
sources: string[]
}
function applyConfidenceThreshold(result: ResolutionResult, threshold = 0.6) {
if (result.confidence < threshold) {
return { ...result, outcome: 'UNKNOWN' }
}
return result
}
The confidence gate is the whole product
The threshold is where the repo stops being a chatbot wrapper and becomes a system design statement. A model can browse, reason, and summarize, but the service layer still gets the final veto. That matters because it converts uncertainty from a hidden flaw into an explicit state.
In practice, that means a market can stay open, a human can inspect the evidence trail, and the resolver can try again later instead of forcing a clean answer out of weak evidence. It is a conservative design, but a useful one if your highest priority is not being confidently wrong.
How it compares to UMA-style resolution
The cleanest comparison is not AI versus humans. It is capital-weighted dispute mechanics versus source-backed machine adjudication. UMA-style resolution is built to make challenges economically meaningful, while this repo is built to make the act of deciding cheap, fast, and auditable.
The final dispute resolution is decided by UMA token holders based on token weight, not evidence weight. A single wallet with sufficient tokens can determine outcomes regardless of facts.
| Dimension | UMA-style resolution | market-ai-resolution |
|---|---|---|
| Decision maker | Token-weighted voters in a dispute process | LLM guided by market rules, live web evidence, and structured outputs |
| Evidence requirements | Evidence matters, but capital and incentives shape the final outcome | Evidence is required, captured, and attached to the answer |
| Time to resolution | Often measured in hours or longer | Designed to resolve in seconds |
| Cost to challenge | High enough to discourage casual disputes | The repo removes challenge economics and relies on the confidence gate |
| Audit trail | Social and economic, with the final result emerging from voting dynamics | Explicit reasoning and source lists are part of the artifact |
| Failure mode | Well-funded actors can dominate the outcome | The model can still be wrong, which is why UNKNOWN exists |
| Best fit | High-stakes markets with mature dispute economics | Lower-friction markets where speed and transparency matter more |
What this project gets right, and what it cannot yet prove
This is a strong architecture for a proof of concept. The separation between AI logic, orchestration, provider abstraction, and event delivery suggests a builder who understands backend boundaries, not just prompting tricks. The stack choices are practical, too: Bun for speed, Hono for a small API surface, and Zod for validation.
But the repo is still early. The market provider is in-memory, the system is centralized around a single runtime and API key, and the article of faith is not production evidence. What it proves is narrower and more interesting: if you ask an oracle to refuse to guess, you can build a system that treats uncertainty as an output.