Ai2 @allen_ai · 01 Sep 2026
BenchMIRT can also estimate how a model will perform on Qs it hasn’t answered. For held-out questions, it correctly predicts whether the model will answer correctly 79% of the time compared with 70% for a simpler benchmark-average baseline.
728Views
3Likes
0Reposts
1Replies
0Quotes
0Bookmarks
Is that a lot?
1.13×vs this author's median645 views is typical
61Percentile for this authorof 18 recent posts
0.25×vs 10K–100K median2 968 views is typical
0.85%Reachviews ÷ followers
0.55%Engagement rateof viewers reacted