Anthropic @AnthropicAI · 01 Sep 2026
New research: Training a Misaligned Reward Seeker What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, we trained an https://t.co/QeXS2Jof3p
648 203Views
2 999Likes
265Reposts
263Replies
176Quotes
1 221Bookmarks
Is that a lot?
0.93×vs this author's median697 298 views is typical
25Percentile for this authorof 4 recent posts
10.0×vs 1M–10M median64 838 views is typical
39.89%Reachviews ÷ followers
0.57%Engagement rateof viewers reacted