tweetindex

Baseten @baseten · 01 Sep 2026

Introducing PACT: a benchmark for whether enterprise AI assistants keep following workplace rules when breaking them is the convenient option. We tested 24 models. The two best are open-weight, and one of them is 27B. Leaderboard, paper, and code: https://t.co/6lhIzKGacS https://t.co/30HzqhXmLD
5 752Views
23Likes
1Reposts
2Replies
2Quotes
0Bookmarks

Is that a lot?

0.57×vs this author's median10 075 views is typical
38Percentile for this authorof 8 recent posts
1.99×vs 10K–100K median2 896 views is typical
30.45%Reachviews ÷ followers
0.49%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →