tweetindex

Jacob Andreas @jacobandreas Β· 01 Jul 2026

New Paper πŸ“„: LMs just want to explain themselves! When we SFT an LM on explanations of its own behaviors, do they learn to actually introspect, or do they merely imitate the original training distribution? We find evidence for the former. Despite training on a static set of https://t.co/AnznpunrJ6
36 438Views
196Likes
36Reposts
6Replies
8Quotes
0Bookmarks

Is that a lot?

3.05Γ—vs this author's median11 940 views is typical
82Percentile for this authorof 11 recent posts
26.3Γ—vs 10K–100K median1 384 views is typical
146.23%Reachviews Γ· followers
0.68%Engagement rateof viewers reacted

Compare with the benchmark table β†’

Open on X β†’