EleutherAI @AiEleuther · 06 Jun 2025
Can you train a performant language models without using unlicensed text? We are thrilled to announce the Common Pile v0.1, an 8TB dataset of openly licensed and public domain text. We train 7B models for 1T and 2T tokens and match the performance similar models like LLaMA 1&2 https://t.co/wHQ4cquqlo
182 657Views
642Likes
157Reposts
22Replies
33Quotes
292Bookmarks
Is that a lot?
40.0×vs this author's median4 567 views is typical
100Percentile for this authorof 7 recent posts
70.8×vs 10K–100K median2 578 views is typical
6.4× audienceReachviews ÷ followers
0.47%Engagement rateof viewers reacted