tweetindex

t11s @transmissions11 · 27 Aug 2026

prefill tok/$ ≈ FLOPs/$ decode tok/$ ≈ HBM BW/$ but not always! consider a chip with poor interconnect pipeline parallelism can rescue prefill, but decode needs many seqs resident to fill a deep pipeline, eating KV capacity and shrinking per-chip batch size + throughput!
12 857Views
89Likes
5Reposts
12Replies
0Quotes
58Bookmarks

Is that a lot?

0.18×vs this author's median72 062 views is typical
0Percentile for this authorof 5 recent posts
4.44×vs 10K–100K median2 896 views is typical
13.68%Reachviews ÷ followers
0.82%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →