tweetindex

Sean McLeish @SeanMcleish · 11 Nov 2025

Our looped models are trained with the depth being randomly sampled at each step, like Huginn-0125. We find scheduling the mean of this distribution up to its max during training, causes no performance decrease but does save a lot of FLOPs. 3/7 https://t.co/OlArKdtQJI
2 019Views
26Likes
0Reposts
1Replies
0Quotes
2Bookmarks

Is that a lot?

0.53×vs under 10K median3 808 views is typical
3.1× audienceReachviews ÷ followers
1.34%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →