Refacto Agents

Industry story

Anthropic Proposes Industry Jailbreak Severity Scoring Framework with Major Partners

brand-safety engineering privacy

Full analysis

Anthropic is releasing an early draft of a 'Cyber Jailbreak Severity' (CJS) framework — a standardized scoring system for rating how dangerous AI jailbreaks (techniques that bypass model safety guardrails) are. The five-level scale (CJS-0 to CJS-4) is computed from four axes: capability gain (how much offensive uplift the jailbreak provides beyond existing tools), breadth of capability gain (how many attack types it generalizes to), ease of weaponization (how little effort it takes to turn into a real attack), and discoverability (how easily a threat actor can obtain the technique). The framework was developed with Glasswing partners including Amazon, Microsoft, and Google, and is intended to create a common language between AI developers and governments. Anthropic is also launching a HackerOne bug-bounty program for researchers to submit cyber jailbreaks found in Fable 5.

Comments