tomaarsen @tomaarsen · 18 Aug 2026
Training works like the other model types: 4 new losses (incl. cached & distillation variants), 5 new evaluators, and a Trainer that takes the same arguments you already know. From a bare ModernBERT-base: 0.1338 -> 0.4831 NanoBEIR mean nDCG@10 in ~25 min on one RTX 3090. https://t.co/P3hANLKMy1
489Views
16Likes
0Reposts
1Replies
0Quotes
5Bookmarks
Is that a lot?
0.77×vs this author's median637 views is typical
27Percentile for this authorof 22 recent posts
0.16×vs under 10K median3 068 views is typical
10.16%Reachviews ÷ followers
3.48%Engagement rateof viewers reacted