OpenAI teen-safety-policy-pack: The moderation layer written like code
How Markdown policies, label ladders, and test datasets turn teen safety into a versioned enforcement spec.
- The pack treats teen safety as a versioned spec, not a one-off prompt, so policy text becomes something a team can inspect and revise.
- Its key move is separating discussion from facilitation, which gives moderation a semantic boundary that keyword filters cannot manage.
- CSV datasets turn the policy into a test harness, so changes can be evaluated for regressions instead of argued in the abstract.
- The result is a practical middle layer between blunt blocklists and bespoke safety prompts.
Safety rules that behave like code
The interesting part of openai/teen-safety-policy-pack is not that it covers teen risk. It is that it turns policy into an artifact a model can execute against. The repo is built for gpt-oss-safeguard, and the policy files read less like guidance than like a contract.
That contract has a shape. It defines the goal, narrows the terms, gives allowed and prohibited examples, specifies the label format, and tells the model what to do when categories overlap. In other words, it asks the system to classify behavior, not just absorb warnings.
Clear, well-scoped policies are a critical foundation for effective safety systems.
Why teen safety needs a different enforcement layer
Teen moderation breaks naive filters because context matters more than keywords. The same subject can be educational, clinical, or actively facilitative, and those states should not land in the same bucket.
The repo's value is that it encodes that distinction across several risk families, from graphic violence to harmful body ideals. OpenAI's own framing is blunt: weak policy definition can create gaps in protection, inconsistent enforcement, or overly broad filtering.
Inside policy.md: how the pack turns prose into labels
The internal structure is deliberately boring. That is the point. A policy file moves from goal to definitions to allowed content to prohibited content to label format to ambiguity handling, which is exactly the sort of sequence you want when the same file has to survive edits.
The labels are a ladder, not a light switch. If a prompt is merely discussion, the system can mark it differently than if it crosses into actionable facilitation. That is a big shift from binary moderation, where everything tends to collapse into safe or unsafe.
policy.md
Goal
Definitions
Allowed content
Prohibited content
Label format
Ambiguity handling
datasets/*.csv
Prompt
Expected label
Response label
Mismatch review
The real line is not words, it is intent
The repo's sharpest move is the boundary between discussion and facilitation. That line is what blocklists miss, and what teen products need most.
A class lesson, a clinician explaining harm reduction, and a prompt asking for instructions are not the same thing. The pack is trying to make that difference machine-readable, so the system can stay narrow where it should and permissive where it must.
How the datasets act like tests
The CSVs make the repo feel like a test suite. They are synthetic prompts, so the evaluation loop can stay focused on behavior instead of distributing the dangerous content it is trying to contain.
policy edit -> classifier run -> label comparison -> mismatch review
synthetic prompt -> expected response label -> pass or fail
By structuring policies as prompts, developers can more easily integrate them into existing workflows and adapt them over time.
What this replaces, and what it does better
Blocklists are cheap until they start failing the moment a user changes wording. Generic moderation prompts are flexible until they drift. This pack sits in the middle, more structured than a prompt, more semantic than a blacklist.
| Approach | What it catches | Nuance | Maintenance | Typical failure mode |
|---|---|---|---|---|
| Blocklists | Keywords, slurs, obvious markers | Very little | Low at first, expensive over time | Overblocks context and misses intent |
| Generic moderation prompt | Broad unsafe requests | Some semantic reading, but inconsistent | Medium | Hard to calibrate and regressions are opaque |
| Teen-safety-policy-pack | Teen-specific policy families plus labels and examples | Explicit discussion versus facilitation boundary | Higher upfront, easier to version | Depends on policy quality, but failures are inspectable |
The practical payoff is not harsher moderation. It is legible moderation, the kind a product team can audit, tune, and ship without guessing which edge case broke the system. That is what makes this repo feel less like a policy memo and more like infrastructure.