tweetindex

Poolside @poolsideai · 17 Aug 2026

Agentic evals are messy. A benchmark score tells you something about model performance, but it also reflects the whole system around it: the harness, sandbox, dependencies, timeouts and sometimes a loophole the agent found in the task. That’s why trajectories matter so much to
8 093Views
65Likes
8Reposts
2Replies
1Quotes
16Bookmarks

Is that a lot?

0.78×vs this author's median10 395 views is typical
22Percentile for this authorof 9 recent posts
4.94×vs 10K–100K median1 637 views is typical
54.75%Reachviews ÷ followers
0.94%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →