tweetindex

Ai2

@allen_ai · Seattle, WA · joined 04 Sep 2015

Breakthrough AI to solve the world's biggest problems. › Join us: https://t.co/MjUpZpKPXJ › Newsletter: https://t.co/k9gGznstwj

86 056Followers
447Following
3 890Posts total
167.7KViews on collected posts

Against accounts of the same size

18 posts from the last 90 days, next to the 10K–100K follower range. below its peers on both reach and engagement.

Median views645this account2 968median for 10K–100K
Reach, %0.75%this account9.18%median for 10K–100K
Engagement, %0.56%this account1.42%median for 10K–100K
MetricThis accountMedian for 10K–100KRatio
Median views per post6452 9680.22×
Reach (views ÷ followers)0.75%9.18%0.08×
Engagement rate0.56%1.42%0.39×

Others in this range →   Compare with another account →   How these benchmarks are built →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

43227 Aug
545
667
45
3.3K1 Sep
14.7K
177
623
428
138
1.6K
728
809
63 Sep

Last 14 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.46%27 Aug
0.37%
0.00%
4.44%
0.57%1 Sep
0.50%
1.69%
0.80%
0.70%
2.17%
0.30%
0.55%
0.87%
0.00%3 Sep

Reactions — likes, reposts, replies and quotes — divided by views. Median for 10K–100K accounts is 1.42%.

What the audience does

Likes57.1%178 in total
Reposts8.0%25 in total
Replies10.9%34 in total
Quotes4.8%15 in total
Bookmarks19.2%60 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Latest posts

@allen_ai BenchMIRT showing BBQ tracks reasoning over safety is a useful eval audit 6 views · 0 likes · 0 reposts · 0 replies 03 Sep 2026 BenchMIRT gives researchers a clearer view of what benchmarks really measure—and could help build evals that are smaller, more focused, & easier to interpret. We’re releasing it openly so others can build on it: 💻 https://t.co/5lnoifXMMK 📄 https://t.co/BkSXP5W5ds 809 views · 7 likes · 0 reposts · 0 replies 01 Sep 2026 BenchMIRT can also estimate how a model will perform on Qs it hasn’t answered. For held-out questions, it correctly predicts whether the model will answer correctly 79% of the time compared with 70% for a simpler benchmark-average baseline. 728 views · 3 likes · 0 reposts · 1 replies 01 Sep 2026 We used BenchMIRT to audit popular LLM evals—and found some quirks. HarmBench mostly tests safety behavior, but its copyright questions depend more on reasoning. XSTest draws on both reasoning & safety, while ToxiGen provides little signal on either. https://t.co/OoGh4OTa 1.6K views · 2 likes · 0 reposts · 2 replies 01 Sep 2026 We applied BenchMIRT across the 16 benchmarks it was trained on to see whether we could make evals more efficient by removing less informative questions. We found keeping just the strongest 10% of Qs preserves nearly the same picture of model strengths as using the full set. 138 views · 2 likes · 0 reposts · 1 replies 01 Sep 2026 We trained BenchMIRT on results from 100 LLMs across 16 benchmarks & 34K+ questions. We didn’t tell it which evals measured what. Two dominant dimensions consistently emerged: general reasoning + safety. https://t.co/GNMJkqAiVq 428 views · 2 likes · 0 reposts · 1 replies 01 Sep 2026 BenchMIRT builds on Item Response Theory (IRT), a technique from psychometrics for measuring abilities from patterns of test responses. The idea: not every question tells you the same amount. Some are harder; some better distinguish stronger models from weaker ones. 623 views · 4 likes · 0 reposts · 1 replies 01 Sep 2026 BenchMIRT works at multiple levels: ◙ For models, it estimates strength on the capabilities reflected in the benchmark set. ◙ For questions, it estimates difficulty & how strongly each Q distinguishes models along those capabilities. 177 views · 2 likes · 0 reposts · 1 replies 01 Sep 2026 Do LLM safety & capability evals measure what they claim to? We built BenchMIRT to audit them + see which model abilities their Qs actually test. On BBQ, a social-bias eval, it found the Qs distinguished models more by reasoning ability than safety. 🧵 https://t.co/BQ8WPTu 14.7K views · 55 likes · 9 reposts · 4 replies 01 Sep 2026 At an event on August 27, we brought together AI researchers, scientists, & medical practitioners to explore what AI needs to do better to meaningfully advance science. Five ideas kept coming up. 🧵 https://t.co/vmqcsXq1EZ 3.3K views · 14 likes · 2 reposts · 2 replies 01 Sep 2026 @allen_ai @ProvSwedish Is it me or every AI2 post seems so impactful? 45 views · 2 likes · 0 reposts · 0 replies 27 Aug 2026 @ProvSwedish This collaboration highlights our approach to AI for science: not replacing scientific judgment, but working alongside scientists to expand what they can explore. Providence Swedish will now deploy AutoDiscovery on additional protected cancer data. https://t.co/lJvg 667 views · 0 likes · 0 reposts · 0 replies 27 Aug 2026 @ProvSwedish The signal involved invasive lobular carcinoma (ILC), a subtype affecting ~48K Americans each year. ILC has long been considered “immune cold,” with relatively low immune activity. AutoDiscovery found more activity than expected. 545 views · 1 likes · 0 reposts · 1 replies 27 Aug 2026 @ProvSwedish Providence Swedish researchers ran AutoDiscovery over The Cancer Genome Atlas (TCGA), a resource built from donated cancer patient samples. AutoDiscovery generated + tested hypotheses across the data, searching for results that challenged expectations. 432 views · 1 likes · 0 reposts · 1 replies 27 Aug 2026 @ProvSwedish One was the ILC immune signal. The researchers confirmed the same pattern in separate patient data. Lab analysis subsequently verified it, finding T-cells around ILC tumors. 123 views · 0 likes · 0 reposts · 1 replies 27 Aug 2026 @ProvSwedish Together, the results suggest ILC warrants broader study for immunotherapy. Read more in the researchers’ scientific report: https://t.co/Kt6Ig1rW0u https://t.co/BePGP621V8 710 views · 1 likes · 0 reposts · 1 replies 27 Aug 2026 AI’s biggest role in science may not be answering questions. It may be helping scientists find which questions are worth asking. At @ProvSwedish, AutoDiscovery surfaced an unexpected immune signal in a heavily studied cancer dataset—and follow-up research confirmed it. 🧵 https:/ 135.8K views · 49 likes · 8 reposts · 11 replies 27 Aug 2026 A Thai research team adapted our Dolma data-curation toolkit to build Mangosteen, a 47B-token corpus for Thai LLMs. They used Dolma to filter widely used web datasets into a smaller corpus that improved Thai LLM performance despite using less data. 🧵 https://t.co/exvf02cy4U htt 6.8K views · 33 likes · 6 reposts · 6 replies 26 Aug 2026

Similar accounts