openai/gpt-oss-safeguard: The Moderation Model That Reads the Rulebook First

A new safety system treats policy as input, not an afterthought. That makes moderation more transparent, more adaptable, and much harder to mistake for a fixed classifier.

9 min read • View on GitHub • More from openai

A giant open rulebook sits on a clean desk while several streams of posts and messages flow toward it like paper ribbons. A verdict stamp appears only after the pages have been read, which explains the article's central idea that policy is a first-class input.
The surprise is not that the model flags content. It is that it reads the rulebook before it decides.
Key Takeaways

Moderation, but the policy comes first

Most moderation systems start with an answer. This one starts with a rulebook. gpt-oss-safeguard reads a policy document and the content being judged at the same time, then reasons to a label instead of forcing every decision through a baked-in filter.

That subtle shift changes the product. OpenAI is not just shipping another safety classifier. It is shipping a way to make policy a runtime input, so a platform can swap rules without retraining the model or rewriting the stack.

A split illustration shows a sealed black box on the left that emits a single label, while a transparent machine on the right exposes policy pages, highlighted clauses, and a visible decision trail. The contrast explains why this model is about inspectable moderation rather than opaque classification.
The old model hides the rulebook inside the weights. This one keeps the rulebook on the table.

Why this is a different kind of safety system

The appeal is not only flexibility. It is governance. A forum, a marketplace, and a game community can all use the same base model, then attach different standards for spam, harassment, cheating, or fake reviews. The rule changes live outside the weights, which makes them easier to audit and easier to revisit.

That matters because the repo is not trying to replace every moderation tool. It is trying to replace the rigid part of the stack with something more adaptable. The smaller 20B model is aimed at lighter deployment, while the larger 120B model gives teams more headroom when policy traffic is heavy.

How the model thinks through a policy

The mechanism is the story. The model takes two inputs, policy text and content, and emits a structured answer in Harmony format. That matters because the reasoning path is visible enough to inspect, while the final label still stays machine-readable.

An interactive map of how policy and content combine into a verdict, with the rulebook treated as a live input.

The spam example makes the idea concrete. The policy separates depiction from request, so a post describing spam is not treated the same as a prompt asking for spam. If the case is fuzzy, the policy tells the model to move one severity level down, which pushes the system toward conservative but explainable decisions.

A close-up shows a five-rung engraved ladder with increasingly dense hatch marks, while a small message card is nudged one rung lower by a careful hand. The image explains the model's severity ladder and the conservative downgrade rule for unclear cases.
The most interesting part is not the label. It is the rule that decides how uncertainty gets resolved.

What changes when policy is not baked in

SystemPolicy sourceOutputAuditabilityCustomizationBest fit
gpt-oss-safeguardDeveloper-written policy at inference timeLabel plus reasoning trailHighHighTeams that need policy control and reviewable decisions
Llama GuardPolicy baked into weightsLabelMediumLow to mediumOpen-weight moderation with a fixed policy set
GPT-4o moderationFixed OpenAI moderation policyLabel and category scoresMediumLowManaged moderation inside OpenAI's stack
Custom fine-tuned classifierTraining data and labelsLabelLow to mediumMediumNarrow domains with stable rules
Keyword filtersHandwritten rulesMatch or blockLowHigh but brittleSimple guardrails and obvious spam

This is why the repo sits in a narrow but important slot. Keyword filters are cheap and blunt. Custom classifiers are tuned but rigid. gpt-oss-safeguard is slower and more operationally demanding than both, but it gives trust and safety teams a living rulebook instead of a frozen label space.

That trade-off makes sense when the policy itself is the product. If you need the same moderation logic to adapt across communities, languages, or business models, a policy reader is more useful than a one-shot classifier. If you need raw throughput and very low cost, the simpler tools still win.