Ai2 @allen_ai · 01 Sep 2026
We applied BenchMIRT across the 16 benchmarks it was trained on to see whether we could make evals more efficient by removing less informative questions. We found keeping just the strongest 10% of Qs preserves nearly the same picture of model strengths as using the full set.
138Views
2Likes
0Reposts
1Replies
0Quotes
0Bookmarks
Is that a lot?
0.21×vs this author's median645 views is typical
17Percentile for this authorof 18 recent posts
0.05×vs 10K–100K median2 968 views is typical
0.16%Reachviews ÷ followers
2.17%Engagement rateof viewers reacted