tweetindex

Snorkel AI @SnorkelAI · 20 Aug 2026

Pretraining data is often curated by quality scores or semantic similarity, but neither necessarily tells us whether selected docs give a model distinct learning signals. In a small GPT-2 test, we took 10 FineWeb documents spanning different FineWeb-Edu quality scores and asked: https://t.co/wu5cjUFhoQ
219Views
1Likes
0Reposts
1Replies
1Quotes
0Bookmarks

Is that a lot?

0.61×vs this author's median359 views is typical
33Percentile for this authorof 9 recent posts
0.13×vs 10K–100K median1 637 views is typical
1.23%Reachviews ÷ followers
1.37%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →