tweetindex

Ai2 @allen_ai · 01 Sep 2026

We applied BenchMIRT across the 16 benchmarks it was trained on to see whether we could make evals more efficient by removing less informative questions. We found keeping just the strongest 10% of Qs preserves nearly the same picture of model strengths as using the full set.
138Views
2Likes
0Reposts
1Replies
0Quotes
0Bookmarks

Is that a lot?

0.21×vs this author's median645 views is typical
17Percentile for this authorof 18 recent posts
0.05×vs 10K–100K median2 968 views is typical
0.16%Reachviews ÷ followers
2.17%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →