tweetindex

John Yang

@jyangballin

CS PhD @Stanford. Created @SWEbench (multi-lingual/modal); SWE-agent; SWE-smith; InterCode; CodeClash; ProgramBench

7 072Followers
1 098Following
968Posts total
869KViews on collected posts

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

751.7K5 May
20.2K
21.8K
17.4K
12.3K
12.6K
12.2K
10.3K
9.4K
1.2K

Last 10 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.27%5 May
0.58%
0.56%
0.71%
0.64%
0.69%
0.91%
0.96%
0.73%
2.78%

Reactions — likes, reposts, replies and quotes — divided by views.

What the audience does

Likes64.6%2 367 in total
Reposts7.2%265 in total
Replies3.7%135 in total
Quotes3.3%120 in total
Bookmarks21.2%778 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Latest posts

@jyangballin @KLieret I first wanted to say this is harder than AGI because no human could do it alone. Then I checked and found out SQLite, FFmpeg and PHP were each written by one developer, back when there was no Stack Overflow to ask. 1.2K views · 30 likes · 1 reposts · 2 replies 05 May 2026 Code: https://t.co/T51AB8DmcI We’ve open sourced the evaluation code, so you can run your agent + model combination on ProgramBench tasks today! Opening submissions for leaderboard, tasks, and tests soon. https://t.co/U6UjUb3xJx 9.4K views · 65 likes · 2 reposts · 2 replies 05 May 2026 ProgramBench is a joint effort across Meta FAIR, Meta TBD, Stanford, Harvard @KLieret (co-first author) @18jeffreyma @parth007_96 @dpedch @sten_sootla @micmylin @pengchengyin @magpie_rayhou @syhw @Diyi_Yang @OfirPress Paper: https://t.co/tCaaFurAFU https://t.co/sSsXDPnjyH 10.3K views · 90 likes · 6 reposts · 3 replies 05 May 2026 ProgramBench is very hard, but it’s solvable by design. While the official best score is 0%, our extended results reveal varying levels of meaningful progress across tasks. We added a lot of interactive plots & tables to our website. https://t.co/S3N2YQh4G6 12.2K views · 104 likes · 2 reposts · 4 replies 05 May 2026 Our 200 instances range from simple (jq, ripgrep) to extremely difficult (ffmpeg, sqlite, compilers). Stats and some team favorites in the images. Improving performance on ProgramBench means that models can increasingly build software from requirements alone. https://t.co/XDDmxb 12.6K views · 84 likes · 1 reposts · 2 replies 05 May 2026 Check out our paper for a bunch of analyses - Models write Python solutions even though originals are in C++/Rust/Go - 98% of runs, models submit before hitting time/turn limits - Given internet access, models Googled solutions 36% of the time, even when told not to https://t.co 12.3K views · 73 likes · 3 reposts · 2 replies 05 May 2026 To make a ProgramBench task: 1. Find a GitHub repo that builds a program 2. Generate 1000s of tests (assert ./program [input] == [output]) Key difference vs. SWE-bench: Our tests are *behavioral*. We don't use unit tests b/c they constrain implementation https://t.co/yCZSI0diHG 17.4K views · 116 likes · 2 reposts · 2 replies 05 May 2026 Several key constraints: - Agent ONLY gets a target executable + some usage doc. - Agent chooses language, designs abstractions, architects the entire program. - No internet access or any other way of cheating. - Executable cannot be decompiled/inspected. https://t.co/qvpHVWv4Xx 21.8K views · 108 likes · 2 reposts · 8 replies 05 May 2026 There's been a wave of case studies on agents building whole programs from scratch, but they test a few projects with hand-tuned setups. ProgramBench formalizes this setting w/ 200 tasks and doubles down on testing, cheat prevention, and task diversity https://t.co/KC5WxwCVEW h 20.2K views · 114 likes · 0 reposts · 3 replies 05 May 2026 How much of SQLite, FFmpeg, PHP compiler can LMs code from scratch? Given just an executable and no starter code or internet access. Introducing ProgramBench: 200 rigorous, whole-repo generation tasks where models design, build, and ship a working program end to end. 🧵 https://t 751.7K views · 1.6K likes · 246 reposts · 107 replies 05 May 2026

Similar accounts