Andrew White 🐦⬛ @andrewwhite01 · 26 Aug 2026
Bixbench3: Can models reproduce the entire analysis from papers? We had agents repro papers end-to-end, with over 1B tokens spent on some. Frontier agents still score below 50%, and due to real problems like environment misconfig, quitting, or faking data 1/6 https://t.co/AYa6pDXYFD
23 057Views
296Likes
43Reposts
15Replies
4Quotes
133Bookmarks
Is that a lot?
1.00×vs this author's median23 057 views is typical
43Percentile for this authorof 7 recent posts
8.38×vs 10K–100K median2 751 views is typical
68.70%Reachviews ÷ followers
1.55%Engagement rateof viewers reacted