tweetindex

Ryan Marten

@ryan_marten

Building @harborframework and @terminalbench

3 270Followers
2 253Following
686Posts total
405.4KViews on collected posts

Against accounts of the same size

16 posts from the last 90 days, next to the under 10K follower range. below its peers on both reach and engagement.

Median views3 162this account2 072median for under 10K
Reach, %96.68%this account153.47%median for under 10K
Engagement, %1.11%this account1.63%median for under 10K
MetricThis accountMedian for under 10KRatio
Median views per post3 1622 0721.53×
Reach (views ÷ followers)96.68%153.47%0.63×
Engagement rate1.11%1.63%0.68×

Others in this range →   Compare with another account →   How these benchmarks are built →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

9.9K23 Jul
5.2K
4.7K
3.2K
4K
2.6K
3.1K
2.9K
3.3K
2K
2.3K
1.9K
1.8K
1.7K

Last 14 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.72%23 Jul
1.12%
1.15%
1.21%
1.15%
1.10%
1.17%
1.18%
1.51%
0.74%
1.46%
0.67%
0.71%
0.76%

Reactions — likes, reposts, replies and quotes — divided by views. Median for under 10K accounts is 1.63%.

What the audience does

Likes70.3%1 632 in total
Reposts5.1%119 in total
Replies5.1%119 in total
Quotes3.7%85 in total
Bookmarks15.8%368 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Latest posts

@modal_labs @AnthropicAI @OpenAI @Google @scale_AI @SnorkelAI @turingcom @gNucleusAI @joinHandshake @ellamindAI @giansegato @boolean_ai Thank you to @modal for providing the sandboxes for developing Frontier-Bench! This was particularly useful for unlocking GPU tasks (e.g. agent 1.7K views · 12 likes · 0 reposts · 1 replies 23 Jul 2026 · Open on X →
@modal_labs @AnthropicAI @OpenAI @Google @scale_AI @SnorkelAI @turingcom @gNucleusAI @joinHandshake @ellamindAI @giansegato Thanks as well to @boolean_ai (missed the tag above). It was a pleasure working together! 1.8K views · 12 likes · 0 reposts · 1 replies 23 Jul 2026 · Open on X →
@modal_labs @AnthropicAI @OpenAI @Google @scale_AI @SnorkelAI @turingcom @gNucleusAI @joinHandshake @ellamindAI Another thank you to @giansegato as an advisor to this project - making sure measurement in benchmarks is rigorous as possible! 1.9K views · 11 likes · 1 reposts · 1 replies 23 Jul 2026 · Open on X →
@ryan_marten Curious the score with Grok 4.5 with Grok Build harness. 2.3K views · 34 likes · 0 reposts · 0 replies 23 Jul 2026 · Open on X →
@modal_labs @AnthropicAI @OpenAI @Google @scale_AI @SnorkelAI @turingcom @gNucleusAI @joinHandshake Additional thank you to another partner of ours who was instrumental for this release. @ellamindAI led the expansion of tasks, via creation and reviews, into the business operation 2K views · 12 likes · 0 reposts · 2 replies 23 Jul 2026 · Open on X →
Frontier-Bench was made possible by our sponsors. Thank you to our compute sponsors @modal_labs, @AnthropicAI, @OpenAI, @Google and our data partners @scale_AI, @SnorkelAI Open Benchmarks, @turingcom, @gNucleusAI, Boolean AI, and @joinHandshake. Frontier-Bench is hosted by 3.3K views · 42 likes · 5 reposts · 2 replies 23 Jul 2026 · Open on X →
@frontierbench is built by the @terminalbench team and @harborframework community Tasks are the beating heart of any worthy benchmark. Thank you to all the task creators and reviewers for all their hard work making this first release. We are lucky enough to have too many 2.9K views · 33 likes · 0 reposts · 1 replies 23 Jul 2026 · Open on X →
Frontier-Bench is a community effort from 100+ task contributors and reviewers. It is the hardest and highest quality benchmark we could make together in the open, but it is just the start. Let's find issues, fix them, and add new tasks together. Let's make this benchmark 3.1K views · 35 likes · 0 reposts · 1 replies 23 Jul 2026 · Open on X →
Special thanks to our senior reviewers @neversupervised, @bla1990so, @TommasoCerruti, @StevenDillmann, @ryne_wang, @dwahdany, @AllenHa13152844, @krauth, advisors @Mike_A_Merrill and Nicholas Carlini and co-leads @alexgshaw, @andykonwinski, @lschmidt3. 2.6K views · 27 likes · 1 reposts · 1 replies 23 Jul 2026 · Open on X →
Agents approach problem solving on Frontier-Bench differently, even at similar pass rates. Fable 5 (33.8%) and Opus 4.8 (21.1%) use more tokens and take fewer actions. GPT-5.6 Sol (34.4%) and Terra (20.8%) use less tokens and take more actions. Overall, GPT-5.6 Sol is ~40% htt 4K views · 43 likes · 0 reposts · 3 replies 23 Jul 2026 · Open on X →
The Frontier-Bench team is already planning upcoming releases. In our next major release, we plan to add new tasks that measure new capabilities, add greater task diversity, and tune timeouts. Tasks are versioned so corresponding trials can be re-used, re-graded, or re-run to h 3.2K views · 37 likes · 1 reposts · 1 replies 23 Jul 2026 · Open on X →
Frontier-Bench discriminates better between frontier and sub-frontier models than Terminal-Bench 2.1. For example, Fable 5 in Claude Code and Opus 4.8 in Claude Code are separated by only 4.9% on Terminal-Bench v2.1, but 12.7% on Frontier-Bench v0.1. https://t.co/gsYh89JRpt 4.7K views · 51 likes · 0 reposts · 2 replies 23 Jul 2026 · Open on X →
Tasks in Frontier-Bench represent domains like software, ML, science, operations, security, hardware, and media. Every task separates the agent container and verifier container. Artifacts from the agent container are downloaded at the end of the trial, logged, and uploaded to ht 5.2K views · 57 likes · 0 reposts · 1 replies 23 Jul 2026 · Open on X →
Frontier-Bench is a continuous benchmark. Today we are releasing v0.1 with 74 difficult, diverse, and high quality tasks. We will be adding new tasks and improving existing ones in regular releases. Benchmarks have historically depreciated over time. We’re rolling out a batch of 9.9K views · 65 likes · 0 reposts · 5 replies 23 Jul 2026 · Open on X →
We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34% h 315.7K views · 1K likes · 97 reposts · 90 replies 23 Jul 2026 · Open on X →
https://t.co/ZVmVZVJpCX 41K views · 143 likes · 14 reposts · 7 replies 23 Jul 2026 · Open on X →

Similar accounts