Ai2 @allen_ai · 01 Sep 2026
We trained BenchMIRT on results from 100 LLMs across 16 benchmarks & 34K+ questions. We didn’t tell it which evals measured what. Two dominant dimensions consistently emerged: general reasoning + safety. https://t.co/GNMJkqAiVq
428Views
2Likes
0Reposts
1Replies
0Quotes
0Bookmarks
Is that a lot?
0.66×vs this author's median645 views is typical
28Percentile for this authorof 18 recent posts
0.14×vs 10K–100K median2 968 views is typical
0.50%Reachviews ÷ followers
0.70%Engagement rateof viewers reacted