tweetindex

Martian @withmartian · 03 Sep 2026

Our insights: ✦ The best LLM system is often not the best LLM. ✦ Advertised token costs are bad predictors of real LLM cost. ✦ LLMs have variable reliability in consistently solving a problem. ✦ Single-model benchmarks are systematically understating current LLM capabilities.
1 718Views
29Likes
0Reposts
5Replies
0Quotes
1Bookmarks

Is that a lot?

1.99×vs this author's median863 views is typical
62Percentile for this authorof 8 recent posts
0.25×vs under 10K median6 872 views is typical
45.90%Reachviews ÷ followers
1.98%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →