tweetindex

Matei Zaharia @matei_zaharia · 03 Sep 2026

Qwen3.8-Flash-Next now runs on a single RTX 5090 at 68.3 tok/s — no extreme quantization, no speculative decoding. · Production-level checkpoint: GB300-validated NVFP4 checkpoint from RadixArk · 63GB host RAM — less than what 1-bit quants of this model need · The 51GB n-gram https://t.co/cxW99MvO4V
53 260Views
463Likes
37Reposts
36Replies
10Quotes
0Bookmarks

Is that a lot?

3.55×vs this author's median15 000 views is typical
80Percentile for this authorof 15 recent posts
40.1×vs 10K–100K median1 328 views is typical
102.30%Reachviews ÷ followers
1.02%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →