tweetindex

Sam Rodriques @SGRodriques · 26 Aug 2026

BixBench3 is out! Our latest eval for biology capabilities in language models. Essentially requires agents to reproduce all work needed to write a paper from raw data. OpenAI is in the lead with 5.6-Sol at 48%, Anthropic is in 4th with Opus 4.8, after Kimi and GLM. Opus 5 has https://t.co/EcMm0cyjnI
8 976Views
167Likes
25Reposts
8Replies
3Quotes
63Bookmarks

Is that a lot?

0.65×vs this author's median13 890 views is typical
0Percentile for this authorof 4 recent posts
3.26×vs 10K–100K median2 751 views is typical
41.80%Reachviews ÷ followers
2.26%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →