Ai2 @allen_ai · 01 Sep 2026
We used BenchMIRT to audit popular LLM evals—and found some quirks. HarmBench mostly tests safety behavior, but its copyright questions depend more on reasoning. XSTest draws on both reasoning & safety, while ToxiGen provides little signal on either. https://t.co/OoGh4OTanG
1 642Views
2Likes
0Reposts
2Replies
1Quotes
0Bookmarks
Is that a lot?
2.55×vs this author's median645 views is typical
72Percentile for this authorof 18 recent posts
0.55×vs 10K–100K median2 968 views is typical
1.91%Reachviews ÷ followers
0.30%Engagement rateof viewers reacted