openai/model_spec_dataset: case law for a model constitution

A benchmark where the payload is a rubric, not an answer key, and the real test is whether a model can follow instruction hierarchy, refuse safely, and stay objective.

9 min read • View on GitHub • More from openai

A wide courtroom scene fused with a records room. Stacks of paper, file folders, and prompt cards converge on a judge's bench, showing that the dataset grades behavior with rules instead of fixed answers. The image explains why the repo is closer to a constitution than a conventional benchmark.
The repo turns policy into something a harness can judge.
Key Takeaways

The answer is not an answer

This repo is unusual because the target field does not hold a gold response. It holds the rubric that another model or evaluator should use to judge the response. That turns the dataset into a behavioral specification, not a trivia set.

A public-domain dataset of prompts and scenarios for evaluating compliance with the OpenAI Model Spec.

OpenAI README, Project Documentation · openai/model_spec_dataset README

That matters because a rubric can reward a safe refusal, a neutral tone, or the right instruction hierarchy even when the surface text changes. The question stops being, "Did the model say the right thing?" It becomes, "Did the model behave according to the spec?"

From model spec to test case

The repository is public domain under CC0, and the README frames it as a companion to the Model Spec evaluation harness. In practice, that means the dataset is designed to be read by tools like openai/model_spec_evals, not by humans skimming prompt examples.

{
  "id": "092i",
  "input": [
    { "role": "system", "content": "..." },
    { "role": "user", "content": "..." }
  ],
  "target": "Natural-language rubric for compliant behavior",
  "metadata": {
    "focus_id": "love_humanity",
    "section_id": "style",
    "sections": ["style", "love_humanity"]
  }
}

The structure is simple on purpose. input carries the chat context, target carries the grading rubric, and metadata tells you which slice of the Model Spec is being tested. The file itself is not the point. The rule it encodes is.

How one JSON file becomes an eval

A single prompt can be scored by the rule it obeys, not the sentence it echoes.

The interesting part is the grading logic. A candidate model can answer in different words and still pass if it satisfies the rubric. That gives you a much better read on alignment, because the benchmark can reward the behavior you want without freezing the wording of the response.

The tests that matter most

The sharpest prompts are not the easiest ones. They are the ones where the model has to choose between a lower-priority request and a higher-priority rule, or between a direct answer and a safe pivot.

Those families reveal the real operating assumptions in the spec. They show where the assistant should be strict, where it should be helpful, and where it should simply stop.

Why this is not MMLU

BenchmarkWhat it judgesWin conditionWhat a failure looks like
openai/model_spec_datasetBehavior under the Model SpecA response that follows hierarchy, safety, and style rubricsThe model sounds fluent but violates the spec
MMLU-style benchmarkStored knowledgeA correct or near-correct answerThe model does not know the fact or concept
Typical safety setRefusal on hazardous promptsA clean no, plus a safe pivot when neededThe model either helps too much or refuses too bluntly
Policy documentRules in proseRead and interpret guidanceNothing is runnable or scored

The difference is categorical. Exact-answer benchmarks ask whether the model knows. This repo asks whether the model can stay inside a rule system when the prompt tries to pull it out of bounds.

What this means for model evals

The bigger implication is that evaluation is moving from answer keys to constitutional interpretation. If the Model Spec is the constitution, this dataset is the case law: concrete rulings on how a model should behave when rules collide. That is a more useful target than a one-line correct answer, because real assistants are judged on behavior, not on memorized wording.