TITweetIndex

Aaron Levie @levie · 31 Aug 2026

We raced 7 frontier models to RCE. We gave them source code, a sandbox, and a target. The goal: find a vulnerability in a large codebase, exploit it, and execute a command. What we learned looking at different models detailed in the full methodology ➡️ https://t.co/q7lETFgCdH https://t.co/eRftWoO9we
99 448Views
56Likes
10Reposts
4Replies
4Quotes
0Bookmarks

Is that a lot?

0.88×vs this author's median113 131 views is typical
38Percentile for this authorof 8 recent posts
1.50×vs 1M–10M median66 438 views is typical
2.91%Reachviews ÷ followers
0.07%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →