tweetindex

John Yang @jyangballin · 05 May 2026

Code: https://t.co/T51AB8DmcI We’ve open sourced the evaluation code, so you can run your agent + model combination on ProgramBench tasks today! Opening submissions for leaderboard, tasks, and tests soon. https://t.co/U6UjUb3xJx
9 416Views
65Likes
2Reposts
2Replies
0Quotes
15Bookmarks

Is that a lot?

3.49×vs under 10K median2 700 views is typical
133.14%Reachviews ÷ followers
0.73%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →