tweetindex

Ryan Greenblatt

@RyanGreenblatt · joined 22 Sep 2023

Chief scientist at Redwood Research (@redwood_ai), focused on technical AI safety research to reduce risks from rogue AIs

19 315Followers
10Following
2 236Posts total
7.5MViews on collected posts

Against accounts of the same size

10 posts from the last 90 days, next to the 10K–100K follower range. shown to more people than peers of the same size.

Median views41 417this account2 896median for 10K–100K
Reach, %214.43%this account9.51%median for 10K–100K
Engagement, %1.29%this account1.56%median for 10K–100K
MetricThis accountMedian for 10K–100KRatio
Median views per post41 4172 89614.3×
Reach (views ÷ followers)2.1× audience9.51%22.6×
Engagement rate1.29%1.56%0.83×

Others in this range →   Compare with another account →   How these benchmarks are built →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

4.1M26 Aug
1.9M
143.2K2 Sep
13.3K
1.2M
59K
21.2K
22K
23.9K
6.6K

Last 10 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.18%26 Aug
0.44%
1.37%2 Sep
1.54%
0.54%
0.77%
1.51%
1.95%
1.21%
1.42%

Reactions — likes, reposts, replies and quotes — divided by views. Median for 10K–100K accounts is 1.56%.

What the audience does

Likes54.8%21 035 in total
Reposts7.4%2 846 in total
Replies2.2%857 in total
Quotes3.2%1 231 in total
Bookmarks32.4%12 420 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Latest posts

I agree with almost all of this and think the point being made here is important. (I'm probably less optimistic about having good mind reading techniques in a few years and I tend to think that AIs might get harder to mind read over time which adds an additional difficulty.) 6.6K views · 90 likes · 1 reposts · 3 replies 02 Sep 2026 @tszzl I think it's unlikely that white-box techniques will be able to provide as much monitorability as CoT currently provides within a year. Currently, they are capable of catching some unverbalized thoughts / plans, but nowhere near as reliably as CoT seems to. I'm optimistic 23.9K views · 249 likes · 22 reposts · 9 replies 02 Sep 2026 I don't think there will be strong+working mech interp in <1 year. Model internals methods could pareto dominate cot monitors in a year via cot losing ~all value, but this isn't exactly encouraging. Depending on AIs decoding other AI's opaque activations for oversight is spo 22K views · 391 likes · 21 reposts · 11 replies 02 Sep 2026 Transparency about the opaque serial depth is great, but this statement is consistent with Astra having a configurable "dial" that is currently set to a low depth but could be trivially increased. We need more info to see how concerning these architectural changes are, 21.2K views · 273 likes · 33 reposts · 12 replies 02 Sep 2026 *monitor-ability* is the invariant that must be preserved, and the real solution will be via strong mechanistic interpretability. i predict in the next year there’ll be mechinterp monitoring that’s pareto optimal to cot monitors 59K views · 389 likes · 13 reposts · 34 replies 02 Sep 2026 I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. OpenAI has worked to preserve and utilize chain-of-thought monitoring since 1.2M views · 5.7K likes · 469 reposts · 236 replies 02 Sep 2026 It's possible OpenAI and others will use these architectures with only a few recurrent iterations (because that's most performant or due to safety concerns). If so, this would only make things moderately worse. But I expect much more opaque reasoning than this in the future... 13.3K views · 184 likes · 11 reposts · 8 replies 02 Sep 2026 OpenAI's newest AI, Astra, is reported to use an 'opaque reasoning' architecture where more of the reasoning occurs in activations instead of natural language. This may be the single worst development for AI security/safety to date. The details of Astra aren't publicly known, 143.2K views · 1.6K likes · 247 reposts · 98 replies 02 Sep 2026 I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'. I semi-jokingly called our efforts a "slop-vestigation" because 1.9M views · 6.5K likes · 1K reposts · 291 replies 26 Aug 2026 METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with l 4.1M views · 5.7K likes · 992 reposts · 155 replies 26 Aug 2026

Similar accounts