Industry story
Anthropic Publishes Detailed Cyber Safeguard Categories for Fable 5
brand-safety engineering privacy
Full analysis
Anthropic has re-deployed Claude Fable 5 globally and released detailed documentation of the cybersecurity safety classifiers — AI systems that detect and block dangerous uses — accompanying the model. The classifiers sort requests into four tiers: Prohibited use (blocked outright, e.g., ransomware, malware development, C2 infrastructure), High-risk dual use (also blocked for now, e.g., penetration testing, exploit development), Low-risk dual use (monitored, sometimes blocked as a 'safety margin'), and Benign use (allowed, e.g., secure coding, log analysis, incident response). Anthropic notes the safety margin is deliberately larger for Fable 5 than for prior models, meaning some legitimate requests will be blocked as a precaution against jailbreaks circumventing higher-risk controls.
Comments