tweetindex

EleutherAI @AiEleuther · 06 Jun 2025

Can you train a performant language models without using unlicensed text? We are thrilled to announce the Common Pile v0.1, an 8TB dataset of openly licensed and public domain text. We train 7B models for 1T and 2T tokens and match the performance similar models like LLaMA 1&2 https://t.co/wHQ4cquqlo
182 657Views
642Likes
157Reposts
22Replies
33Quotes
292Bookmarks

Is that a lot?

40.0×vs this author's median4 567 views is typical
100Percentile for this authorof 7 recent posts
70.8×vs 10K–100K median2 578 views is typical
6.4× audienceReachviews ÷ followers
0.47%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →