alphaXiv @askalphaxiv · 12 May 2026
Reinforcing Recursive Language Models Can a 4B model learn to recursively call itself to answer hard long-context questions? We RL fine-tuned a small model to behave as a native RLM. On evidence selection across scientific papers, our 4B RLM matches Sonnet 4.6 in quality https://t.co/COd3rVn1bE
98 240Views
596Likes
83Reposts
15Replies
10Quotes
498Bookmarks
Is that a lot?
36.6×vs this author's median2 682 views is typical
100Percentile for this authorof 4 recent posts
35.7×vs 10K–100K median2 751 views is typical
182.17%Reachviews ÷ followers
0.72%Engagement rateof viewers reacted