tweetindex

Adam Karvonen @a_karvonen · 21 Aug 2026

We also use the investigations as training data, teaching models to explain their own behavior. This generalizes to held-out OOD evals, such as detecting when a hint changed an answer, though results vary with training format (more on this later).
1 037Views
14Likes
1Reposts
1Replies
0Quotes
0Bookmarks

Is that a lot?

0.85×vs this author's median1 223 views is typical
27Percentile for this authorof 11 recent posts
0.17×vs under 10K median6 244 views is typical
23.98%Reachviews ÷ followers
1.54%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →