GPT OSS Safeguard 20B
gpt-oss-safeguard is a first open weight reasoning model specifically trained for safety classification tasks to help classify text content based on customizable policies. As a fine-tuned version of gpt-oss, gpt-oss-safeguard is designed to follow explicit written policies that you provide. This enables bring-your-own-policy Trust & Safety AI, where your own taxonomy, definitions, and thresholds guide classification decisions. Well crafted policies unlock gpt-oss-safeguard's reasoning capabilities, enabling it to handle nuanced content, explain borderline decisions, and adapt to contextual factors.
- Reasoning
Added Feb 23, 2026
Model weightsPricing
Auto routing · per 1M tokens- Input
- $0.075
- Output
- $0.30
Specifications
- Context window
- 128K
- Max output
- 16.4K
- Parameters
- 20B / 3.6B
- Total / active
- Avg output (7d)
- 192 tokens
- Longer than 14% of models
Benchmarks
No public benchmark scores for this model yet.
Providers
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…