tweetindex

John Yang @jyangballin · 05 May 2026

There's been a wave of case studies on agents building whole programs from scratch, but they test a few projects with hand-tuned setups. ProgramBench formalizes this setting w/ 200 tasks and doubles down on testing, cheat prevention, and task diversity https://t.co/KC5WxwCVEW https://t.co/JdhhFPWhUs
20 154Views
114Likes
0Reposts
3Replies
0Quotes
22Bookmarks

Is that a lot?

7.46×vs under 10K median2 700 views is typical
2.8× audienceReachviews ÷ followers
0.58%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →