elie @eliebakouch · 02 Sep 2026
a few thoughts on "recurrent depth" transformers the main question: why recurrent depth instead of just scaling depth? it's not faster at inference or training* since you still go through the full "effective depth", the advantage is storage (for kv cache storage btw you could do https://t.co/SD29JZECtU
105 847Views
743Likes
48Reposts
39Replies
12Quotes
0Bookmarks
Is that a lot?
3.36×vs this author's median31 519 views is typical
71Percentile for this authorof 7 recent posts
38.5×vs 10K–100K median2 751 views is typical
4.4× audienceReachviews ÷ followers
0.80%Engagement rateof viewers reacted