Industry story
Anthropic Proposes Industry Jailbreak Severity Scoring Framework with Major Partners
brand-safety engineering privacy
Anthropic is trying to do for AI jailbreaks what CVSS did for software vulnerabilities — give the industry a shared severity language before the chaos compounds. The Cyber Jailbreak Severity framework runs a five-level scale across four axes: capability gain, breadth, weaponization ease, and discoverability. Amazon, Microsoft, and Google are already on board, which matters more than the framework itself — a scoring system nobody uses is just a paper. The HackerOne bug-bounty tied to Claude 3.5 is the real test; researchers submitting real jailbreaks against a live model will stress-test whether CJS holds up in practice.
Full analysis
Anthropic is releasing an early draft of a 'Cyber Jailbreak Severity' (CJS) framework — a standardized scoring system for rating how dangerous AI jailbreaks (techniques that bypass model safety guardrails) are. The five-level scale (CJS-0 to CJS-4) is computed from four axes: capability gain (how much offensive uplift the jailbreak provides beyond existing tools), breadth of capability gain (how many attack types it generalizes to), ease of weaponization (how little effort it takes to turn into a real attack), and discoverability (how easily a threat actor can obtain the technique). The framework was developed with Glasswing partners including Amazon, Microsoft, and Google, and is intended to create a common language between AI developers and governments. Anthropic is also launching a HackerOne bug-bounty program for researchers to submit cyber jailbreaks found in Fable 5.
Comments