tweetindex

shirish @shiri_shh · 03 Sep 2026

We got 46% fewer errors than the single best LLM across the 16 most used benchmarks (TerminalBench, LiveCodeBench, etc). Here's how that's possible and what each model can achieve when used optimally (every benchmarks misses the majority of model capabilities) 👇 Interactive https://t.co/OXOwcwLCQR
171 016Views
157Likes
15Reposts
47Replies
52Quotes
0Bookmarks

Is that a lot?

3.09×vs this author's median55 421 views is typical
56Percentile for this authorof 9 recent posts
50.5×vs 10K–100K median3 383 views is typical
3.9× audienceReachviews ÷ followers
0.16%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →