tweetindex

Ofir Press

@OfirPress · NYC · joined 25 Jun 2016

I push the AI frontier by building tough benchmarks with amazing people. SWE-bench, SWE-agent, SciCode, AlgoTune. Postdoc @Princeton. PhD @nlpnoah @UW.

19 490Followers
9 279Following
3 434Posts total
1.1MViews on collected posts

Against accounts of the same size

11 posts from the last 90 days, next to the 10K–100K follower range. shown widely, but few of those viewers react.

Median views3 207this account3 598median for 10K–100K
Reach, %16.45%this account11.59%median for 10K–100K
Engagement, %0.59%this account1.38%median for 10K–100K
MetricThis accountMedian for 10K–100KRatio
Median views per post3 2073 5980.89×
Reach (views ÷ followers)16.45%11.59%1.42×
Engagement rate0.59%1.38%0.42×

Others in this range →   Compare with another account →   How these benchmarks are built →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

751.1K5 May
142.1K
13K26 Jun
146.8K
2.5K
3.1K
19921 Jul
3.2K27 Aug
3.1K
4.7K28 Aug
22.8K1 Sep
2.9K
6.3K

Last 13 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.27%5 May
0.80%
0.37%26 Jun
0.16%
0.20%
0.22%
3.52%21 Jul
0.09%27 Aug
0.59%
0.94%28 Aug
1.27%1 Sep
1.20%
1.36%

Reactions — likes, reposts, replies and quotes — divided by views. Median for 10K–100K accounts is 1.38%.

What the audience does

Likes74.0%3 276 in total
Reposts8.3%368 in total
Replies4.5%198 in total
Quotes2.8%122 in total
Bookmarks10.4%462 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Latest posts

SWE-bench Multimodal was used in today's Anthropic launch. No model passes the 60% mark. Lots of room for growth https://t.co/M66pS8deNm 6.3K views · 75 likes · 4 reposts · 7 replies 01 Sep 2026 We made SWE-bench Multimodal more than a year ago, but the instances were never public- now they are! 480 new tasks, JavaScript/HTML/CSS. Anthropic has used it in Mythos and Fable, and now you can too. 2.9K views · 27 likes · 3 reposts · 5 replies 01 Sep 2026 Releasing SWE-bench Multimodal v2.0 today 480 tasks where a coding agent must interpret visual assets like screenshots, diagrams, recordings to diagnose and fix a bug in a repository. https://t.co/tEvvRFzwCr 22.8K views · 238 likes · 35 reposts · 15 replies 01 Sep 2026 people sometimes confuse diff length with task hardness. Linus just pushed a bugfix to the kernel, which he claimed took a "debug session from hell" to find. it's 1 line of code. https://t.co/c1JMs3fhZF 4.7K views · 40 likes · 0 reposts · 4 replies 28 Aug 2026 Yes- the gold mine that @closji & @jyangballin discovered in SWE-bench is that if you find a scalable way to mine benchmark tasks, not only will you have an easier time building the benchmark, but you'll also have a straightforward method to getting lots of training examples! 3.1K views · 13 likes · 2 reposts · 3 replies 27 Aug 2026 @OfirPress interesting that you didn't put "scalable" in the tldr; looks like being able to automatically construct tasks is the only way to keep up with model updates. curious to hear your thoughts on that. 3.2K views · 2 likes · 0 reposts · 0 replies 27 Aug 2026 @OfirPress @woj_zaremba @OriolVinyalsML @ilyasut ha, must’ve been some serious second thoughts lol 199 views · 3 likes · 3 reposts · 1 replies 21 Jul 2026 @OfirPress @woj_zaremba @OriolVinyalsML @ilyasut ``` weight tying solved something real. it's kinda wild how long it stuck around even after transformers changed everything else 3.1K views · 4 likes · 3 reposts · 0 replies 26 Jun 2026 @OfirPress @woj_zaremba @OriolVinyalsML @ilyasut 10 years of improving perplexity on a 1M token dataset is wild context for a technique that still gets used 2.5K views · 3 likes · 0 reposts · 2 replies 26 Jun 2026 I started doing language modeling research in Jan 2016, using the LSTM from the "RNN Regularization" model by @woj_zaremba @OriolVinyalsML and @ilyasut, trying to improve its perplexity on Penn Treebank (1M toks). This led to the weight tying method, sometimes still used today. 146.8K views · 214 likes · 9 reposts · 16 replies 26 Jun 2026 It's been ten years since @OfirPress first wrote on how to get started in deep learning -- and most of what he wrote is still relevant today! Congratulations Ofir!👏👏👏 https://t.co/milhUsTdeZ 13K views · 45 likes · 1 reposts · 1 replies 26 Jun 2026 1) Our team at Meta has a tough new coding benchmark challenging models to code entire programs including ffmpeg and the PHP compiler from scratch. 2) Top accuracy is 0% 3) We will be making the benchmark harder. 142.1K views · 1K likes · 62 reposts · 37 replies 05 May 2026 How much of SQLite, FFmpeg, PHP compiler can LMs code from scratch? Given just an executable and no starter code or internet access. Introducing ProgramBench: 200 rigorous, whole-repo generation tasks where models design, build, and ship a working program end to end. 🧵 https://t 751.1K views · 1.6K likes · 246 reposts · 107 replies 05 May 2026

Similar accounts