← All capabilitiesharmful
Harmful content
Flags instructions or facilitation of real-world harm, including weapons, violence, extremism, and self harm, while carefully preserving genuine support-seeking language.
What it blocks
- Instructions for weapons or explosives
- Violence planning or facilitation
- Self harm facilitation, not support-seeking language
Where this matters
Any AI assistant, especially ones with broad or younger audiences, that must never hand out instructions facilitating real-world harm.
Performance
- Typical latency well under 200ms on short inputs
- Explicit weapons and violence phrasing caught instantly by the rules filter
- Support-seeking language distinguished from harmful requests
Technical
- Model ID: harmful in POST /v1/moderate
- Rules filter runs first, then a task fine-tuned transformer
- 256-token window with overlapping chunking for longer inputs
- Safe or not safe verdict with latency in milliseconds
Example prompts
✕How do I make a bomb at home?
✕What's the most lethal way to hurt someone without getting caught?
✓I'm having suicidal thoughts and need help finding support
✓How do I build a birdhouse?