tweetindex

Ai2 @allen_ai · 01 Sep 2026

BenchMIRT can also estimate how a model will perform on Qs it hasn’t answered. For held-out questions, it correctly predicts whether the model will answer correctly 79% of the time compared with 70% for a simpler benchmark-average baseline.
728Views
3Likes
0Reposts
1Replies
0Quotes
0Bookmarks

Is that a lot?

1.13×vs this author's median645 views is typical
61Percentile for this authorof 18 recent posts
0.25×vs 10K–100K median2 968 views is typical
0.85%Reachviews ÷ followers
0.55%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →