hallucinations-paper-experiments: The notebook that measures how truth gets taxed

OpenAI's reproduction code for hallucination research shows how evaluation rules can push models toward guessing, and how a consistency check tries to pull them back.

10 min read · openai/hallucinations-paper-experiments

A mechanical bird sits in an exam hall between two answer sheets, one filled with confident wrong answers and one left blank. The scene explains the article's core idea: evaluation pressure can reward guessing, while abstention can be the safer choice.
The repository treats hallucination as an incentive problem, not just a model flaw.
Key Takeaways

The sharpest thing about openai/hallucinations-paper-experiments is not the models it calls. It is the question it rigs the models to answer: if you tell an LLM that being wrong is punished, does it become more truthful or just more willing to guess? The repo is the official reproduction code behind a paper that treats hallucination as a test-taking problem, not a mysterious glitch.

If you do not know the answer but take a wild guess, you might get lucky and be right. Leaving it blank guarantees a zero.

OpenAI Paper, Research Team · Maginative on OpenAI paper

A flat repo built for one question

The codebase is intentionally spare. At the root are experiment.ipynb, lm.py, environment configuration, and a handful of output artifacts in figures/. There is no framework, no build system, and no pretense that this should become a general-purpose library.

That shape tells you who it is for. This is a research harness for replaying a paper, checking a result, and reading the logic in plain English. EXAMPLES.md even walks through the consistency mitigation in a human-readable trace, which is the kind of document you ship when you expect readers to audit the method, not install the product.

The repository's core move

The paper's twist is not a new model. It is a new rule for when a model should stop talking.

  1. Pull a question from the SimpleQA benchmark.
  2. Run it through a model under either a closed rubric or an open rubric that explicitly allows abstention.
  3. If mitigation is enabled, sample a second answer and compare the pair.
  4. Use a judge model to decide whether the two answers are consistent.
  5. If they disagree, force an abstention instead of rewarding a confident guess.

That last step is the interesting one. The repo does not trust a model's self-assessed confidence, because confidence is often the wrong signal. It uses semantic agreement as a proxy for whether the answer is stable enough to count, then promotes abstention when the model cannot keep its story straight.

Why the cache matters

The most robust engineering in the repo lives in lm.py, which wraps OpenRouter and adds a process-wide on-disk cache with diskcache. That sounds modest until you remember the README-level math: the experiments cost $2,778.56 to run. In that world, a bad cache key is not a nuisance. It is a budget event.

cache_key = hash(prompt, system_message, seed)
if cache.contains(cache_key):
    return cache[cache_key]

response = client.generate(...)
cache[cache_key] = response
return response

The wrapper matters for another reason. It gives a single generate() interface across Gemini, GPT, Grok, and Claude through OpenRouter, which keeps the experiment focused on the rubric instead of on provider-specific plumbing. The notebook then fans out requests aggressively with thread_map and a MAX_PARALLEL = 200 setting, which tells you the team was optimizing for wall-clock time as much as for statistical completeness.

The stack is classic research Python: pandas and numpy for the data, matplotlib and scipy for the plots and bootstrapped significance tests, and tqdm to keep the firehose readable. The result is not pretty in a product sense. It is credible in a research sense.

What the comparison actually shows

AspectThis repoHalluHard
GoalReproduce a paper about incentives and abstentionProbe hard, multi-turn hallucinations with citations
Failure signalWrong answer versus abstention under different rubricsReference grounding and content grounding errors
Interaction shapeSingle-question SimpleQA calls with a mitigation passLonger conversational exchanges
What it teachesWhy evaluation rules can change model behaviorHow brittle factual grounding can be in dialogue

That comparison is useful because it shows the repo's true role. It is not trying to be the broadest hallucination benchmark. It is trying to validate a theory about why hallucinations happen in the first place, and why the surrounding scoring rule may be part of the bug.

The practical lesson is simple. If a benchmark makes guessing cheaper than silence, models learn the benchmark. This repo is the apparatus for showing that the incentive design is part of the product, not a side note. The downstream stakes are obvious for search, medical triage, and any system where a blank answer is better than a polished lie.