tweetindex
DE

Jacob Austin

@jacobaustin132

I sometimes do AI research. I also play piano and climb. NYC. Previously @GoogleDeepMind, @Google Brain. Opinions my own

7 626Followers
940Following
575Posts total
529.1KViews on collected posts

Neueste Beiträge

@jacobaustin132 Jocab, quickly glanced through it. What a wonderful book! Thank you. Can you please add on computation (memory & FLOP) for LORA to this chapter: All the Transformer Math You Need to Know https://t.co/HroJtFGsDp 769 views · 5 likes · 0 reposts · 0 replies Open on X →
We want this to be a living book, so please ask questions and give us feedback. We'll continue adding to it as time goes on. Without further ado, here’s a link to the beginning: https://t.co/UHYnso8xD1 11/11 4.6K views · 54 likes · 1 reposts · 1 replies Open on X →
The rest of the book is a set of practical guides: how to write and profile parallel JAX code, and how to apply the previous two sections to real models like LLaMA-3. We also have worked problems at the end of each section if you like homework: https://t.co/WgPXed3K1h 8/n https:/
4.1K views · 34 likes · 0 reposts · 1 replies Open on X →
This book was co-written with @_sholtodouglas, @charliexychen, @pchoy95, @albertwebson, @vinayramasesh, @froystig, @anselmlevskaya, @sharadvikram, and Fede Lebron, building on prior ideas by @reinerp and @jekbradbury. 10/n 6.9K views · 50 likes · 0 reposts · 1 replies Open on X →
Now for the good stuff! You may have heard of data or tensor parallelism, FSDP or pipelining. But why choose one over the other? Short answer: each adds communication, and the one with the lowest cost depends on the model. Part 5 dives into this: https://t.co/QiPyAb0HLr 6/n https
6.7K views · 57 likes · 3 reposts · 1 replies Open on X →
Now that we’ve talked about training, we need to talk about serving. How expensive should a model be to serve? What kind of latency can we expect? What are prefill and generation? How do we build an efficient inference service? We talk about this here: https://t.co/W0y0Z5G3qX 7/n
4.5K views · 38 likes · 0 reposts · 1 replies Open on X →
5 years ago, there were many ML architectures, but today, there is (mostly) only one. _You should know the Transformer inside and out!_ How many FLOPs or params in LLaMA-3? How expensive is attention vs. a feed-forward block? You'll know after reading https://t.co/RxQz2qKmdM 5/n
6K views · 56 likes · 1 reposts · 1 replies Open on X →
Scaling an LLM involves distributing — a.k.a. "sharding" — its weights across multiple TPUs. To run it, we have to add cross-chip communication. Part 3 describes the TPU's communication primitives, and simple rules for multiplying sharded matrices: https://t.co/uOqxtFWHEw 4/n htt
GIF
6.7K views · 55 likes · 1 reposts · 1 replies Open on X →
The secret is to think in terms of basic system resources — compute, memory, and bandwidth — and calculate which one limits our performance. From this we can estimate the cost, runtime, and optimal parallelism strategy for any given LLM: https://t.co/UHYnso7ZNt 2/n https://t.co/N
12K views · 120 likes · 8 reposts · 2 replies Open on X →
A big chunk of this book is dedicated to understanding the hardware that provides those system resources. We emphasize TPUs in this book, but the principles and math can be adapted to GPUs too. Part 2 explains the TPU in detail: https://t.co/SFgIzEe0JL 3/n https://t.co/pnTP81uOI3
GIF
9.2K views · 66 likes · 4 reposts · 1 replies Open on X →
Making LLMs run efficiently can feel scary, but scaling isn’t magic, it’s math! We wanted to demystify the “systems view” of LLMs and wrote a little textbook called “How To Scale Your Model” which we’re releasing today. 1/n https://t.co/jnb5kTLD5V
467.5K views · 1.9K likes · 387 reposts · 25 replies Open on X →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

467.5K4 Feb
9.2K
12K
6.7K
6K
4.5K
6.7K
6.9K
4.1K
4.6K
7695 Feb

Last 11 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.50%4 Feb
0.77%
1.08%
0.85%
0.97%
0.86%
0.91%
0.74%
0.84%
1.21%
0.65%5 Feb

Reactions — likes, reposts, replies and quotes — divided by views.

What the audience does

Likes43.6%2 450 in total
Reposts7.2%405 in total
Replies0.6%35 in total
Bookmarks48.5%2 724 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Ähnliche Konten