tweetindex
IT

Greg Diamos

@GregoryDiamos

I build AI supercomputers

3 504Followers
142Following
863Posts total
168.7KViews on collected posts

Ultimi post

@intel I guess it is doing a good job? 1.1K views · 5 likes · 0 reposts · 0 replies Open on X →
@intel Agree that this is with basic bf16 and it is possible to make the active param count higher through better weight compression. 1.5K views · 9 likes · 0 reposts · 2 replies Open on X →
It seems like this got a lot of likes. Remember this is an experiment - don't expect it to be a good production model. I'd do a multi trillion token foundation model run and then fine tune it before using it. ~6k token/sec/core would take a couple months on a 32-core CPU. 6.5K views · 79 likes · 0 reposts · 2 replies Open on X →
Fitting active params into the L2 cache gets up to 14,572 tok/sec on an @intel emerald rapids CPU with AMX 7.1K views · 97 likes · 2 reposts · 3 replies Open on X →
Data is doing more of the work than it used to. Every source in our mixture is a curated artifact built with large models Training a model this small on them is distillation When models of this size were last studied seriously such corpora did not exist 9.8K views · 145 likes · 6 reposts · 1 replies Open on X →
The loss does not saturate. Across a 4.91B-token run, smoothed training loss falls monotonically within each curriculum phase and is still descending at the end 10.1K views · 82 likes · 0 reposts · 1 replies Open on X →
The behaviours appear early. We assumed in-context copying, positional manipulation and arithmetic would need budgets well beyond one core’s reach. They do not. At 259M tokens — about nine hours — a block-routed MoE with E = 128 experts -- begins learning. 11.7K views · 141 likes · 0 reposts · 0 replies Open on X →
@GregoryDiamos Wow, this is really cool Gregory! Thanks for sharing. So partly what you're saying is L2 cache is really cool and we all should want more of it! 3.9K views · 16 likes · 0 reposts · 1 replies Open on X →
I think we should revisit outrageously small neural nets. I needed a 10k tok/s CPU model for data processing. So I gave Anthropic claude code a pile of tokens to build one. It made three interesting discoveries: https://t.co/eR28fVZdDI https://t.co/Pv1zDOPLS0
117K views · 2.1K likes · 112 reposts · 85 replies Open on X →

Rispetto ad account della stessa dimensione

9 post degli ultimi 90 giorni, accanto alla fascia di under 10K follower. esattamente sulla mediana della sua fascia di follower.

Visualizzazioni mediane7 086questo account2 488mediana per under 10K
Copertura, %202.23%questo account163.70%mediana per under 10K
Interazione, %1.20%questo account1.42%mediana per under 10K
MetricaQuesto accountMediana per under 10KRapporto
Visualizzazioni mediane per post7 0862 4882.85×
Copertura (visualizzazioni ÷ follower)2.0× audience163.70%1.24×
Tasso di interazione1.20%1.42%0.85×

Altri account di questa fascia →   Confronta con un altro account →   Come sono costruiti questi parametri →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

117K7 Sep
3.9K
11.7K
10.1K
9.8K
7.1K
6.5K
1.5K
1.1K

Last 9 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

1.93%7 Sep
0.44%
1.20%
0.82%
1.55%
1.45%
1.25%
0.75%
0.45%

Reactions — likes, reposts, replies and quotes — divided by views. Median for under 10K accounts is 1.42%.

What the audience does

Likes52.2%2 626 in total
Reposts2.4%120 in total
Replies1.9%95 in total
Quotes0.3%16 in total
Bookmarks43.2%2 173 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Account simili