Jacob Andreas @jacobandreas Β· 01 Jul 2026
New Paper π: LMs just want to explain themselves! When we SFT an LM on explanations of its own behaviors, do they learn to actually introspect, or do they merely imitate the original training distribution? We find evidence for the former. Despite training on a static set of https://t.co/AnznpunrJ6
36 438Views
196Likes
36Reposts
6Replies
8Quotes
0Bookmarks
Is that a lot?
3.05Γvs this author's median11 940 views is typical
82Percentile for this authorof 11 recent posts
26.3Γvs 10Kβ100K median1 384 views is typical
146.23%Reachviews Γ· followers
0.68%Engagement rateof viewers reacted