Steven Dillmann @StevenDillmann · 27 Aug 2026
Terminal-Bench-Science discriminates between frontier models as well as Terminal-Bench 3.0, while pushing pass rates down more than 10 points for every model evaluated on both. Opus 5: 43% → 30% GPT-5.6 Sol: 34% → 22% Fable 5: 34% → 21% Opus 4.8: 21% → 11% GPT-5.6 Terra: 21% https://t.co/TM8TnxsYjm
3 188Views
33Likes
2Reposts
1Replies
0Quotes
6Bookmarks
Is that a lot?
1.36×vs this author's median2 351 views is typical
64Percentile for this authorof 14 recent posts
0.84×vs under 10K median3 808 views is typical
2.0× audienceReachviews ÷ followers
1.13%Engagement rateof viewers reacted