tweetindex

Snorkel AI @SnorkelAI · 02 Sep 2026

We evaluated Fable 5.1 against Opus 5 on Terminal-Bench+, our proprietary set of frontier coding tasks. Fable 5.1 leads on debugging (87%) and games (88%), using 58% fewer tokens per successful run. Failure analysis shows the errors cluster in execution, not reasoning. Full https://t.co/ACpVsRFRap
1 473Views
28Likes
7Reposts
8Replies
1Quotes
8Bookmarks

Is that a lot?

4.10×vs this author's median359 views is typical
78Percentile for this authorof 9 recent posts
0.90×vs 10K–100K median1 637 views is typical
8.30%Reachviews ÷ followers
2.99%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →