Center for AI Safety @CAIS · 31 Jul 2026
What does a "fixed empirical threshold" actually look like? A model could be considered unsafe to release if (Capabilities > X) AND [ (Refusal Rate < Y) OR (Jailbreak Success Rate > Z) ] The Virology Capabilities Test or ExploitGym could measure hazardous biological or cyber
1 856Views
17Likes
4Reposts
3Replies
0Quotes
3Bookmarks
Is that a lot?
0.16×vs this author's median11 738 views is typical
14Percentile for this authorof 7 recent posts
0.64×vs 10K–100K median2 896 views is typical
17.13%Reachviews ÷ followers
1.29%Engagement rateof viewers reacted