tweetindex

Adam Karvonen

@a_karvonen

ML Researcher, doing MATS with Owain Evans. I prefer email to DM.

4 325Followers
744Following
1 603Posts total
61.8KViews on collected posts

Against accounts of the same size

11 posts from the last 90 days, next to the under 10K follower range. reaches fewer people than peers of the same size.

Median views1 223this account6 244median for under 10K
Reach, %28.28%this account415.20%median for under 10K
Engagement, %1.16%this account1.06%median for under 10K
MetricThis accountMedian for under 10KRatio
Median views per post1 2236 2440.20×
Reach (views ÷ followers)28.28%4.2× audience0.07×
Engagement rate1.16%1.06%1.09×

Others in this range →   Compare with another account →   How these benchmarks are built →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

39.6K21 Aug
2.1K
6.5K
1.2K
5K
1K
2.8K
1K
1.1K
912
433

Last 11 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.75%21 Aug
1.16%
0.46%
1.39%
0.82%
2.64%
0.97%
1.54%
3.12%
3.07%
0.69%

Reactions — likes, reposts, replies and quotes — divided by views. Median for under 10K accounts is 1.06%.

What the audience does

Likes64.7%472 in total
Reposts5.8%42 in total
Replies2.2%16 in total
Quotes2.2%16 in total
Bookmarks25.2%184 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Latest posts

@a_karvonen super cool stuff! How did you go about collecting a dataset of “in the wild” behaviors? 433 views · 2 likes · 0 reposts · 1 replies 21 Aug 2026 @euan_ong @thesubhashk @saprmarks TL;DR: We make an interp eval on "in the wild" behaviors, but interp tools provide no uplift. @JacobSteinhardt @ChowdhuryNeil @ArthurConmy @sen_r @RyanGreenblatt @banburismus_ @jacobandreas @BronsonSchoen 912 views · 26 likes · 0 reposts · 2 replies 21 Aug 2026 Paper: https://t.co/hNPDIun0N2 Blog post: https://t.co/jCKf5XpL3W Coauthors: @euan_ong, @thesubhashk, and @saprmarks. Work done as part of the Anthropic Fellows Program. 1.1K views · 32 likes · 1 reposts · 1 replies 21 Aug 2026 We also use the investigations as training data, teaching models to explain their own behavior. This generalizes to held-out OOD evals, such as detecting when a hint changed an answer, though results vary with training format (more on this later). 1K views · 14 likes · 1 reposts · 1 replies 21 Aug 2026 The pipeline also surfaces unfaithful CoT in the wild. For example, when asked to pick a show from a list, Qwen3-8B just picks the first show from the list and then makes up a reason to support its choice. https://t.co/U89TmxS0Ah 2.8K views · 24 likes · 1 reposts · 1 replies 21 Aug 2026 Why? Sometimes the tools help, but they also mislead. Their outputs name things you can already see in the transcript, like the cause (the variable name) and the behavior (writing code). But they rarely mention the causal relationship: "the variable name causes the bug". 1K views · 25 likes · 1 reposts · 1 replies 21 Aug 2026 We test three activation-based tools that succeeded in prior auditing games: activation oracles, natural language autoencoders, and SAEs. None beat just reading the transcript. The result is robust to hyperparameters, prompting, and a Fable goal loop attempting to improve it. h 5K views · 33 likes · 2 reposts · 2 replies 21 Aug 2026 We evaluate agents on their ability to diagnose the cause and predict the outcome of prompt edits, such as "renaming `max/min` to `a/b` stops the model from making the coding mistake." Does providing an agent with interp tools help over just reading the transcript? 1.2K views · 15 likes · 1 reposts · 1 replies 21 Aug 2026 The causes vary widely: specific words in the prompt, quirks of the model (a misremembered fact about an actor), or abstract properties (a user's angry tone). https://t.co/L9pGQ4Unof 6.5K views · 26 likes · 1 reposts · 1 replies 21 Aug 2026 The eval data is produced by our pipeline CHIVE (Counterfactual Hypothesis Investigation Via Edits) that can be ran on any model or prompt source. In this example, it finds that Gemma is making a coding error due to misleading variable names. https://t.co/0HhFSLAMu7 2.1K views · 22 likes · 1 reposts · 1 replies 21 Aug 2026 A good explanation of a model's behavior should help you make predictions in related situations. We turn this into an eval, with thousands of real behaviors found in the wild. Can interp tools help here? On average, no. 🧵 https://t.co/yN3t8G4n40 39.6K views · 253 likes · 33 reposts · 4 replies 21 Aug 2026

Similar accounts