John Yang ✓
@jyangballin
CS PhD @Stanford. Created @SWEbench (multi-lingual/modal); SWE-agent; SWE-smith; InterCode; CodeClash; ProgramBench
7 072Followers
1 098Following
968Posts total
869KViews on collected posts
Growth & engagement
How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.
Views per post
751.7K5 May
20.2K
21.8K
17.4K
12.3K
12.6K
12.2K
10.3K
9.4K
1.2K
Last 10 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.
Engagement rate per post
0.27%5 May
0.58%
0.56%
0.71%
0.64%
0.69%
0.91%
0.96%
0.73%
2.78%
Reactions — likes, reposts, replies and quotes — divided by views.
What the audience does
Likes64.6%2 367 in total
Reposts7.2%265 in total
Replies3.7%135 in total
Quotes3.3%120 in total
Bookmarks21.2%778 in total
Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.
The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.
Latest posts
@jyangballin @KLieret I first wanted to say this is harder than AGI because no human could do it alone.
Then I checked and found out SQLite, FFmpeg and PHP were each written by one developer, back when there was no Stack Overflow to ask.
1.2K views · 30 likes
· 1 reposts · 2 replies
05 May 2026
Code: https://t.co/T51AB8DmcI
We’ve open sourced the evaluation code, so you can run your agent + model combination on ProgramBench tasks today! Opening submissions for leaderboard, tasks, and tests soon. https://t.co/U6UjUb3xJx
9.4K views · 65 likes
· 2 reposts · 2 replies
05 May 2026
ProgramBench is a joint effort across Meta FAIR, Meta TBD, Stanford, Harvard
@KLieret (co-first author) @18jeffreyma @parth007_96 @dpedch @sten_sootla @micmylin @pengchengyin @magpie_rayhou @syhw @Diyi_Yang @OfirPress
Paper: https://t.co/tCaaFurAFU https://t.co/sSsXDPnjyH
10.3K views · 90 likes
· 6 reposts · 3 replies
05 May 2026
ProgramBench is very hard, but it’s solvable by design.
While the official best score is 0%, our extended results reveal varying levels of meaningful progress across tasks. We added a lot of interactive plots & tables to our website. https://t.co/S3N2YQh4G6
12.2K views · 104 likes
· 2 reposts · 4 replies
05 May 2026
Our 200 instances range from simple (jq, ripgrep) to extremely difficult (ffmpeg, sqlite, compilers). Stats and some team favorites in the images.
Improving performance on ProgramBench means that models can increasingly build software from requirements alone. https://t.co/XDDmxb
12.6K views · 84 likes
· 1 reposts · 2 replies
05 May 2026
Check out our paper for a bunch of analyses
- Models write Python solutions even though originals are in C++/Rust/Go
- 98% of runs, models submit before hitting time/turn limits
- Given internet access, models Googled solutions 36% of the time, even when told not to https://t.co
12.3K views · 73 likes
· 3 reposts · 2 replies
05 May 2026
To make a ProgramBench task:
1. Find a GitHub repo that builds a program
2. Generate 1000s of tests (assert ./program [input] == [output])
Key difference vs. SWE-bench: Our tests are *behavioral*.
We don't use unit tests b/c they constrain implementation https://t.co/yCZSI0diHG
17.4K views · 116 likes
· 2 reposts · 2 replies
05 May 2026
Several key constraints:
- Agent ONLY gets a target executable + some usage doc.
- Agent chooses language, designs abstractions, architects the entire program.
- No internet access or any other way of cheating.
- Executable cannot be decompiled/inspected. https://t.co/qvpHVWv4Xx
21.8K views · 108 likes
· 2 reposts · 8 replies
05 May 2026
There's been a wave of case studies on agents building whole programs from scratch, but they test a few projects with hand-tuned setups.
ProgramBench formalizes this setting w/ 200 tasks and doubles down on testing, cheat prevention, and task diversity
https://t.co/KC5WxwCVEW h
20.2K views · 114 likes
· 0 reposts · 3 replies
05 May 2026
How much of SQLite, FFmpeg, PHP compiler can LMs code from scratch? Given just an executable and no starter code or internet access.
Introducing ProgramBench: 200 rigorous, whole-repo generation tasks where models design, build, and ship a working program end to end. 🧵 https://t
751.7K views · 1.6K likes
· 246 reposts · 107 replies
05 May 2026
Similar accounts
Gabriel Dechichi ✓
@gdechichi
14.1K followers
· 3.5K posts
Brandon Meier
@BrandonMeier
14.1K followers
· 1.9K posts
John Dolan
@JohnDolanAuthor
14.1K followers
· 209.9K posts
Christian ✓
@SoyChristian
14.1K followers
· 4.1K posts
Union of Jewish Students
@UJS_UK
14.1K followers
· 11.2K posts
João Neves ✓
@Joao_neves87
14.1K followers
· 33 posts
AS Monaco Mercato
@ASM_Mercato
14.1K followers
· 11.7K posts
Kan, The Artwork Guy ✗🇳🇬 ✓
@FindingKan
14.1K followers
· 179.3K posts