Refacto Agents

Industry story

Anthropic Publishes Detailed Cyber Safeguard Categories for Fable 5

brand-safety engineering privacy

Anthropic isn't just releasing a more capable model — it's publishing the rulebook alongside it. The Claude Fable 5 cybersecurity classifiers sort requests into four tiers, with the top two (ransomware, malware, exploit development, C2 infrastructure) blocked outright. The deliberate choice to widen the safety margin means some legitimate penetration testing and security research gets caught in the net — Anthropic is accepting false positives to reduce the risk that jailbreaks can walk the model up to the line and then over it. Security practitioners will feel the friction; that's the point.

Full analysis

Anthropic has re-deployed Claude Fable 5 globally and released detailed documentation of the cybersecurity safety classifiers — AI systems that detect and block dangerous uses — accompanying the model. The classifiers sort requests into four tiers: Prohibited use (blocked outright, e.g., ransomware, malware development, C2 infrastructure), High-risk dual use (also blocked for now, e.g., penetration testing, exploit development), Low-risk dual use (monitored, sometimes blocked as a 'safety margin'), and Benign use (allowed, e.g., secure coding, log analysis, incident response). Anthropic notes the safety margin is deliberately larger for Fable 5 than for prior models, meaning some legitimate requests will be blocked as a precaution against jailbreaks circumventing higher-risk controls.

Comments