tweetindex

Ai2 @allen_ai · 01 Sep 2026

Do LLM safety & capability evals measure what they claim to? We built BenchMIRT to audit them + see which model abilities their Qs actually test. On BBQ, a social-bias eval, it found the Qs distinguished models more by reasoning ability than safety. 🧵 https://t.co/BQ8WPTu5nR https://t.co/r4Bi173STl
14 679Views
55Likes
9Reposts
4Replies
5Quotes
28Bookmarks

Is that a lot?

22.8×vs this author's median645 views is typical
89Percentile for this authorof 18 recent posts
4.95×vs 10K–100K median2 968 views is typical
17.06%Reachviews ÷ followers
0.50%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →