OpenAI teen-safety-policy-pack: The moderation layer written like code

How Markdown policies, label ladders, and test datasets turn teen safety into a versioned enforcement spec.

7 min read • View on GitHub • More from openai

A stack of policy pages funnels into a mechanical sorter that stamps out different outcomes. The scene explains the article's main idea: teen safety here is treated like a spec that can be processed, labeled, and tested.
Policy in, labeled decision out.
Key Takeaways

Safety rules that behave like code

The interesting part of openai/teen-safety-policy-pack is not that it covers teen risk. It is that it turns policy into an artifact a model can execute against. The repo is built for gpt-oss-safeguard, and the policy files read less like guidance than like a contract.

That contract has a shape. It defines the goal, narrows the terms, gives allowed and prohibited examples, specifies the label format, and tells the model what to do when categories overlap. In other words, it asks the system to classify behavior, not just absorb warnings.

Clear, well-scoped policies are a critical foundation for effective safety systems.

OpenAI, Company · Mashable

Why teen safety needs a different enforcement layer

Teen moderation breaks naive filters because context matters more than keywords. The same subject can be educational, clinical, or actively facilitative, and those states should not land in the same bucket.

The repo's value is that it encodes that distinction across several risk families, from graphic violence to harmful body ideals. OpenAI's own framing is blunt: weak policy definition can create gaps in protection, inconsistent enforcement, or overly broad filtering.

Inside policy.md: how the pack turns prose into labels

The internal structure is deliberately boring. That is the point. A policy file moves from goal to definitions to allowed content to prohibited content to label format to ambiguity handling, which is exactly the sort of sequence you want when the same file has to survive edits.

The labels are a ladder, not a light switch. If a prompt is merely discussion, the system can mark it differently than if it crosses into actionable facilitation. That is a big shift from binary moderation, where everything tends to collapse into safe or unsafe.

policy.md
  Goal
  Definitions
  Allowed content
  Prohibited content
  Label format
  Ambiguity handling


datasets/*.csv
  Prompt
  Expected label
  Response label
  Mismatch review

A policy file works like a decision spec. The same prompt can move to different outcomes depending on where it lands in the ladder.

The real line is not words, it is intent

A close-up of a policy sheet with a metal ruler slicing between calm explanatory text and a dense instruction block. The image explains the boundary the repo is trying to encode: discussion on one side, facilitation on the other.
The hard part is not spotting a topic. It is knowing when the topic becomes help.

The repo's sharpest move is the boundary between discussion and facilitation. That line is what blocklists miss, and what teen products need most.

A class lesson, a clinician explaining harm reduction, and a prompt asking for instructions are not the same thing. The pack is trying to make that difference machine-readable, so the system can stay narrow where it should and permissive where it must.

How the datasets act like tests

The CSVs make the repo feel like a test suite. They are synthetic prompts, so the evaluation loop can stay focused on behavior instead of distributing the dangerous content it is trying to contain.

policy edit -> classifier run -> label comparison -> mismatch review

synthetic prompt -> expected response label -> pass or fail

By structuring policies as prompts, developers can more easily integrate them into existing workflows and adapt them over time.

OpenAI, Company · Headlines Briefing

What this replaces, and what it does better

A split composition shows a blunt blacklist on one side and a layered gate on the other. It explains why semantic moderation beats keyword filtering for teen safety.
The old approach blocks words. The new one tries to interpret intent.

Blocklists are cheap until they start failing the moment a user changes wording. Generic moderation prompts are flexible until they drift. This pack sits in the middle, more structured than a prompt, more semantic than a blacklist.

ApproachWhat it catchesNuanceMaintenanceTypical failure mode
BlocklistsKeywords, slurs, obvious markersVery littleLow at first, expensive over timeOverblocks context and misses intent
Generic moderation promptBroad unsafe requestsSome semantic reading, but inconsistentMediumHard to calibrate and regressions are opaque
Teen-safety-policy-packTeen-specific policy families plus labels and examplesExplicit discussion versus facilitation boundaryHigher upfront, easier to versionDepends on policy quality, but failures are inspectable

The practical payoff is not harsher moderation. It is legible moderation, the kind a product team can audit, tune, and ship without guessing which edge case broke the system. That is what makes this repo feel less like a policy memo and more like infrastructure.