Perry Dong @perryadong · 30 Jul 2026
To fix this, our key idea: loosen the Q-function's constraint to base-policy actions We propose Initialization via Policy Ensemble (IPE) — initialize Q-function on rollouts from diverse policies, giving it real coverage so fine-tuning learns the RL optimal Q-function (5/6) https://t.co/v24czDtdHT
1 903Views
9Likes
1Reposts
1Replies
0Quotes
1Bookmarks
Is that a lot?
1.07×vs this author's median1 780 views is typical
50Percentile for this authorof 6 recent posts
0.27×vs under 10K median7 156 views is typical
169.01%Reachviews ÷ followers
0.58%Engagement rateof viewers reacted