tweetindex

AI at Meta @AIatMeta · 20 Aug 2026

Today we’re also previewing WildArtifactBench, an internal evaluation framework designed to assess agents on complex, real-world tasks across diverse deliverable formats. By using win rates and Elo scores from human and agentic preference judges rather than strict ground-truth https://t.co/eXWkTrOmPo
150 469Views
632Likes
59Reposts
92Replies
15Quotes
163Bookmarks

Is that a lot?

2.92×vs this author's median51 501 views is typical
83Percentile for this authorof 12 recent posts
6.87×vs 100K–1M median21 891 views is typical
17.84%Reachviews ÷ followers
0.53%Engagement rateof viewers reacted

Compare with the benchmark table →

Open on X →