jietang @jietang · 20 Aug 2026
An excellent history of scaling laws from @jietang. In 2020, we explored the limits of sparsity in Switch Transformers by routing each token to only 1 out of 2048 experts (in retrospect, a bold choice). The model had fewer than 3B activated parameters, but 1.6T total parameters https://t.co/2CjbN9325T
165 470Views
650Likes
53Reposts
13Replies
10Quotes
0Bookmarks
Is that a lot?
1.01×vs this author's median164 162 views is typical
57Percentile for this authorof 7 recent posts
57.1×vs 10K–100K median2 896 views is typical
2.3× audienceReachviews ÷ followers
0.44%Engagement rateof viewers reacted