gpt-oss-safeguard technical report
OpenAI just released two open-weight reasoning models purpose-built for content moderation at scale.

Why it matters
OpenAI is releasing gpt-oss-safeguard models (120B and 20B parameters) designed to apply custom safety policies to content at scale. This signals a shift toward open-sourcing safety infrastructure—traditionally a closed competitive advantage—and democratizing policy-driven content classification for enterprises.
The key facts
6 to knowTwo model sizes released: gpt-oss-safeguard-120b and gpt-oss-safeguard-20b
Open-weight models (not closed API)
Post-trained from gpt-oss base models for reasoning capability
Designed for policy-driven content labeling
Includes baseline safety evaluations against underlying gpt-oss models
Published October 29, 2025
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are two open-weight reasoning models post-trained from the gpt-oss models and trained to reason from a provided policy in order to label content under that policy. In this report, we describe gpt-oss-safeguard’s capabilities and provide our baseline…