tweetindex

Steven Dillmann

@StevenDillmann

AI4Science PhD @Stanford @StanfordAILab Terminal-Bench-Science Lead @harborframework Research Intern @allen_ai Prev. @Cambridge_Uni, @NASAJPL, @imperialcollege

1 593Followers
1 770Following
384Posts total
246KViews on collected posts

Against accounts of the same size

14 posts from the last 90 days, next to the under 10K follower range. reaches fewer people than peers of the same size.

Median views2 351this account3 808median for under 10K
Reach, %147.58%this account249.49%median for under 10K
Engagement, %1.03%this account1.30%median for under 10K
MetricThis accountMedian for under 10KRatio
Median views per post2 3513 8080.62×
Reach (views ÷ followers)147.58%2.5× audience0.59×
Engagement rate1.03%1.30%0.79×

Others in this range →   Compare with another account →   How these benchmarks are built →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

214.4K27 Aug
4.9K
3.7K
3.7K
3.2K
2.1K
2.3K
2.4K
2K
2.5K
2K
948
1.2K
789

Last 14 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.47%27 Aug
0.86%
0.87%
0.93%
1.13%
1.31%
1.30%
2.05%
1.42%
1.27%
1.88%
0.63%
0.84%
0.89%

Reactions — likes, reposts, replies and quotes — divided by views. Median for under 10K accounts is 1.30%.

What the audience does

Likes60.7%1 065 in total
Reposts9.1%159 in total
Replies4.3%75 in total
Quotes4.2%73 in total
Bookmarks21.8%383 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

Followers by day

4 Sep

Daily snapshots since 04 Sep 2026; the dashed line is the starting count.

Latest posts

Extra special thanks to @AllenHa13152844 - a driving force behind this project from the start. 789 views · 7 likes · 0 reposts · 0 replies 27 Aug 2026 Also huge thank you to @bespokelabsai for supporting our leaderboard experiments - couldn't have released it in time without you! 1.2K views · 8 likes · 1 reposts · 1 replies 27 Aug 2026 @StevenDillmann This is really awesome work! Will be so great to have a high quality, scientifically meaningful benchmark for the community to hillclimb Soon we'll start doing for science broadly what the models have already started doing for math 948 views · 4 likes · 1 reposts · 1 replies 27 Aug 2026 Leaderboard, announcement, and full contributor list now live on our website. https://t.co/GeamJxOPfp 11/11 🚀 2K views · 30 likes · 2 reposts · 5 replies 27 Aug 2026 Special thanks to our amazing reviewer team @AllenHa13152844, @neversupervised, @joejans70565123, @brownight7453, @stebo85, @s_aganezov, @BenBlaiszik, @Krauth, @Robertljg, @hou_han_, @Andrew_Y_Liu, @mars_gao_, @0xrobertzhang, our scientific advisors @sarameghanbeery, Jo Dunkley, 2.5K views · 27 likes · 2 reposts · 1 replies 27 Aug 2026 Terminal-Bench-Science was made possible by our research partners @StanfordAILab, @stanfordnlp, @StanfordHAI, @2077AI, @MLFoundations, @AllenInstitute, @allen_ai, and our compute sponsors @AnthropicAI, @Google, @SnorkelAI, @UniPat_AI, @modal, @Kimi_Moonshot, @SpaceXAI, @Zai_org. 2K views · 26 likes · 1 reposts · 1 replies 27 Aug 2026 We're already working on Terminal-Bench-Science 0.2. Deadline is October 5. If you're a practicing scientists and want to contribute a task to our next release, join us and reach out. https://t.co/DJj5GIj2T6 10/n 2.4K views · 42 likes · 4 reposts · 2 replies 27 Aug 2026 Performance also varies by scientific domain. @AnthropicAI and @OpenAI's best models dominate every domain, except for the engineering sciences, where @SpaceXAI's Grok 4.6 ties GPT-5.6 Sol for second (14.8%) at lower cost and token usage. Opus 5 takes the top spot in every htt 2.3K views · 26 likes · 3 reposts · 1 replies 27 Aug 2026 Terminal-Bench-Science is built by the @terminalbench and @harborframework team in collaboration with scientists at @Stanford, @MIT, @Princeton, @Caltech, @Harvard, @UCBerkeley, @UTAustin, @BU_Tweets, @GeorgiaTech, @UMich, @SLAClab, @argonne, @imperialcollege, @UniofOxford, 2.1K views · 25 likes · 2 reposts · 1 replies 27 Aug 2026 Terminal-Bench-Science discriminates between frontier models as well as Terminal-Bench 3.0, while pushing pass rates down more than 10 points for every model evaluated on both. Opus 5: 43% → 30% GPT-5.6 Sol: 34% → 22% Fable 5: 34% → 21% Opus 4.8: 21% → 11% GPT-5.6 Terra: 21% htt 3.2K views · 33 likes · 2 reposts · 1 replies 27 Aug 2026 Cost Pareto frontier: Low-cost end: GPT-5.6 Luna, Kimi K3, GPT-5.6 Terra High-performance end: GPT-5.6 Sol, Claude Opus 5 GPT-5.6 Sol matches the performance of Claude Fable 5 at less than 1/3 of the cost ($4.2k vs $14.2k) Token Pareto frontier: Low-token end: Kimi K3 https:/ 3.7K views · 28 likes · 2 reposts · 2 replies 27 Aug 2026 Tasks are contributed by researchers worldwide through an open review process. Every task goes through three layers of review: a domain reviewer, a technical reviewer, and a bar raiser. 920 proposals → 464 approved → 386 PRs opened → 70 tasks in Terminal-Bench-Science 0.1. https 3.7K views · 29 likes · 2 reposts · 1 replies 27 Aug 2026 Terminal-Bench-Science 0.1 has 70 tasks across the life, physical, earth, mathematical, and engineering sciences. Each one is contributed by a researcher porting their own workflows into challenging agentic tasks, on problems they actually care about. 2/n https://t.co/7RcegRHWXs 4.9K views · 39 likes · 2 reposts · 1 replies 27 Aug 2026 We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains. An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions https: 214.4K views · 741 likes · 135 reposts · 57 replies 27 Aug 2026

Similar accounts