tweetindex

John Schulman @johnschulman2 · 04 Sep 2026

Can a model learn to explain its behaviors, such as why it ignored a user request or made a coding error? We trained models on thousands of explanations of their own in-the-wild behaviors. Training on this single general dataset shows generalization to held-out evals. 🧵 https://t.co/muiBtPbQRM
57 570Views
204Likes
20Reposts
4Replies
5Quotes
0Bookmarks

Is that a lot?

0.56×vs this author's median102 925 views is typical
17Percentile for this authorof 6 recent posts
37.4×vs 10K–100K median1 540 views is typical
71.93%Reachviews ÷ followers
0.41%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →