tweetindex

Ian Leslie @mrianleslie · 29 Aug 2026

This seems like a substantive, unsolvable problem in this paradigm: It is really hard to RL the right behaviour without inadvertently reinforcing covering up and deceiving. If you punish the agentss for deception, you are reinforcing successful deception.
10 799Views
48Likes
5Reposts
2Replies
1Quotes
0Bookmarks

Is that a lot?

0.95×vs this author's median11 335 views is typical
38Percentile for this authorof 8 recent posts
3.19×vs 10K–100K median3 383 views is typical
41.51%Reachviews ÷ followers
0.52%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →