Robert Nishihara @robertnishihara · 01 Sep 2026
How does one RL post-train a 397B model for long-horizon knowledge work? 👩💼 We share every step we took to bring Qwen 3.5 397B from 16.1% Pass@1 to 27.3% on APEX-Agents using DPPO, including final models weights and the full training script🚀 This is the first of many works from https://t.co/VDbGqqzRhQ
129 209Views
709Likes
100Reposts
16Replies
21Quotes
0Bookmarks
Is that a lot?
5.71×vs this author's median22 635 views is typical
71Percentile for this authorof 7 recent posts
47.0×vs 10K–100K median2 751 views is typical
7.5× audienceReachviews ÷ followers
0.66%Engagement rateof viewers reacted