Perry Dong @perryadong · 30 Jul 2026
We break down the core reason: the pretrained critic targets the wrong value function Even in the limit of an optimal pretraining distribution, its action preferences diverge from the optimal RL critic's (3/6) https://t.co/PCkOZVLD26
2 060Views
7Likes
1Reposts
1Replies
0Quotes
1Bookmarks
Is that a lot?
1.16×vs this author's median1 780 views is typical
67Percentile for this authorof 6 recent posts
0.29×vs under 10K median7 156 views is typical
182.95%Reachviews ÷ followers
0.44%Engagement rateof viewers reacted