Five models, each fine-tuned for one slice of moderation. Same request format, same billing, same rules-first pipeline.
| Attribute | Jailbreak | Vulgar | PII | Scam | Harmful |
|---|---|---|---|---|---|
| What it catches | Prompt injection and instruction overrides | Profanity and targeted abuse | Personal information in text | Phishing, fraud, and impersonation | Instructions for real-world harm |
| Where it matters | LLM chat surfaces and copilots | Public comments and live chat | Assistants handling customer data | Messaging and marketplaces | Broad or young audiences |
| Rules filter catches | Known injection phrasing | Exact profanity | Emails, cards, ID formats | Common scam phrasing | Explicit harm phrasing |
| Transformer resolves | Obfuscated and encoded attempts | Context and targeted insults | Contextual data requests | Nuanced scam structures | Support-seeking vs harm |
| Typical latency | <200ms | <200ms | <200ms | <200ms | <200ms |
All five share one endpoint, one billing rate, and the rules-first pipeline. Only the task they are fine-tuned for changes.
Sign up, get 1,000 free requests, and run every model against real text before you decide anything.