tweetindex

Jacob Andreas @jacobandreas · 03 Jul 2026

Higher benchmark scores do not always mean better models for users. Why? We claim that RL teaches LMs to be correct but not how to be correct: code can pass tests but be unreadable; explanations can be right but unclear. How do we train LMs to be right in the right way? (1/n) https://t.co/UsJYA7Q0pp
31 597Views
131Likes
33Reposts
5Replies
4Quotes
0Bookmarks

Is that a lot?

2.65×vs this author's median11 940 views is typical
64Percentile for this authorof 11 recent posts
22.8×vs 10K–100K median1 384 views is typical
126.80%Reachviews ÷ followers
0.55%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →