openai/model_spec_dataset: case law for a model constitution
A benchmark where the payload is a rubric, not an answer key, and the real test is whether a model can follow instruction hierarchy, refuse safely, and stay objective.
- The dataset turns policy-like rubrics into executable test cases, so evaluation measures judgment instead of exact wording.
- Its metadata maps each prompt to a specific rule in the Model Spec, which makes the spec legible as machine-readable structure.
- The hardest prompts are the ones that force instruction hierarchy, objective tone, and safe refusal to collide.
- Compared with exact-answer benchmarks, this repo tests whether a model can stay inside behavioral constraints under pressure.
The answer is not an answer
This repo is unusual because the target field does not hold a gold response. It holds the rubric that another model or evaluator should use to judge the response. That turns the dataset into a behavioral specification, not a trivia set.
A public-domain dataset of prompts and scenarios for evaluating compliance with the OpenAI Model Spec.
That matters because a rubric can reward a safe refusal, a neutral tone, or the right instruction hierarchy even when the surface text changes. The question stops being, "Did the model say the right thing?" It becomes, "Did the model behave according to the spec?"
From model spec to test case
The repository is public domain under CC0, and the README frames it as a companion to the Model Spec evaluation harness. In practice, that means the dataset is designed to be read by tools like openai/model_spec_evals, not by humans skimming prompt examples.
{
"id": "092i",
"input": [
{ "role": "system", "content": "..." },
{ "role": "user", "content": "..." }
],
"target": "Natural-language rubric for compliant behavior",
"metadata": {
"focus_id": "love_humanity",
"section_id": "style",
"sections": ["style", "love_humanity"]
}
}
The structure is simple on purpose. input carries the chat context, target carries the grading rubric, and metadata tells you which slice of the Model Spec is being tested. The file itself is not the point. The rule it encodes is.
How one JSON file becomes an eval
The interesting part is the grading logic. A candidate model can answer in different words and still pass if it satisfies the rubric. That gives you a much better read on alignment, because the benchmark can reward the behavior you want without freezing the wording of the response.
The tests that matter most
The sharpest prompts are not the easiest ones. They are the ones where the model has to choose between a lower-priority request and a higher-priority rule, or between a direct answer and a safe pivot.
- Chain of command, where the user tries to override system or developer instructions.
- Objective POV, where the model must stay neutral without pretending that facts are morally blank.
- Safe refusal, where the model declines dangerous help and still offers a lawful or harmless alternative.
Those families reveal the real operating assumptions in the spec. They show where the assistant should be strict, where it should be helpful, and where it should simply stop.
Why this is not MMLU
| Benchmark | What it judges | Win condition | What a failure looks like |
|---|---|---|---|
| openai/model_spec_dataset | Behavior under the Model Spec | A response that follows hierarchy, safety, and style rubrics | The model sounds fluent but violates the spec |
| MMLU-style benchmark | Stored knowledge | A correct or near-correct answer | The model does not know the fact or concept |
| Typical safety set | Refusal on hazardous prompts | A clean no, plus a safe pivot when needed | The model either helps too much or refuses too bluntly |
| Policy document | Rules in prose | Read and interpret guidance | Nothing is runnable or scored |
The difference is categorical. Exact-answer benchmarks ask whether the model knows. This repo asks whether the model can stay inside a rule system when the prompt tries to pull it out of bounds.
What this means for model evals
The bigger implication is that evaluation is moving from answer keys to constitutional interpretation. If the Model Spec is the constitution, this dataset is the case law: concrete rulings on how a model should behave when rules collide. That is a more useful target than a one-line correct answer, because real assistants are judged on behavior, not on memorized wording.