tweetindex

AlphaSignal @AlphaSignalAI · 18 Jul 2026

What our test is (and is not): We run a coding-agent repair harness (Signaldesk), not a frontend board and not Kimi Code. > 13 planted-bug tasks · 7 models · same rules for all > Network-off Docker sandboxes > Held-out tests only at scoring > Pass = visible + held-out + https://t.co/Nrg7JUbCU1
6 132Views
26Likes
1Reposts
2Replies
0Quotes
3Bookmarks

Is that a lot?

8.50×vs this author's median721 views is typical
57Percentile for this authorof 7 recent posts
4.84×vs 10K–100K median1 266 views is typical
37.02%Reachviews ÷ followers
0.47%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →