tweetindex

Barret Zoph @barret_zoph · 27 Oct 2025

Our latest post explores on-policy distillation, a training approach that unites the error-correcting relevance of RL with the reward density of SFT. When training it for math reasoning and as an internal chat assistant, we find that on-policy distillation can outperform other https://t.co/ltPsdNjajD
1 947 693Views
2 776Likes
402Reposts
59Replies
154Quotes
0Bookmarks

Is that a lot?

673×vs 10K–100K median2 896 views is typical
68.8× audienceReachviews ÷ followers
0.17%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →