tweetindex

AlphaSignal @AlphaSignalAI · 01 Sep 2026

A moderation model can read your policy, react to it, and still enforce it incorrectly. Tested @MistralAI's 3B Shieldstral on 124 policy decisions. The model clearly responds to runtime rules. Every unfamiliar-policy match scored above its unrelated control. But the failures https://t.co/jdw55BM1Vi
721Views
5Likes
0Reposts
2Replies
1Quotes
0Bookmarks

Is that a lot?

1.00×vs this author's median721 views is typical
43Percentile for this authorof 7 recent posts
0.57×vs 10K–100K median1 266 views is typical
4.35%Reachviews ÷ followers
1.11%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →