Dmytro Dzhulgakov @dzhulgakov · 29 Aug 2026
Along the way, we found a few peculiar behaviors. EvalScope prompt included explicit “think step by step” instructions. It caused Flash to significantly overthink (2-3x more tokens). It was still a fair comparison as we used identical eval harness across endpoints AIME and GPQA https://t.co/GpnVq9hm8k
4 176Views
31Likes
0Reposts
1Replies
1Quotes
5Bookmarks
Is that a lot?
0.54×vs this author's median7 671 views is typical
17Percentile for this authorof 6 recent posts
0.74×vs under 10K median5 664 views is typical
56.63%Reachviews ÷ followers
0.79%Engagement rateof viewers reacted