tweetindex

John Yang @jyangballin · 05 May 2026

To make a ProgramBench task: 1. Find a GitHub repo that builds a program 2. Generate 1000s of tests (assert ./program [input] == [output]) Key difference vs. SWE-bench: Our tests are *behavioral*. We don't use unit tests b/c they constrain implementation https://t.co/yCZSI0diHG
17 367Views
116Likes
2Reposts
2Replies
4Quotes
17Bookmarks

Is that a lot?

6.43×vs under 10K median2 700 views is typical
2.5× audienceReachviews ÷ followers
0.71%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →