tweetindex

Einsia @EinsiaAI · 21 Aug 2026

3/ Across 29 model × harness × effort configurations: Scores range from 0.065 to 0.288, with an average of just 0.166. Opus 5 is the best. GPT, Claude, and Kimi stay fairly close overall. Median exploration cost per task jumps from $1.69 to $34.60—but spending more doesn’t https://t.co/Rr1TtNBjr9
799Views
16Likes
0Reposts
2Replies
0Quotes
2Bookmarks

Is that a lot?

1.41×vs this author's median565 views is typical
57Percentile for this authorof 7 recent posts
0.19×vs under 10K median4 164 views is typical
136.58%Reachviews ÷ followers
2.25%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →