tweetindex

Alexander Panfilov

@kotekjedi_ml

MATS 9.0 | PhD @ELLISInst_Tue & @MPI_IS doing AI Safety & Adversarial ML

9 233Followers
390Following
523Posts total
5.1MViews on collected posts

Against accounts of the same size

14 posts from the last 90 days, next to the under 10K follower range. shown widely, but few of those viewers react.

Median views122 599this account2 387median for under 10K
Reach, %1327.83%this account160.90%median for under 10K
Engagement, %0.56%this account1.63%median for under 10K
MetricThis accountMedian for under 10KRatio
Median views per post122 5992 38751.4×
Reach (views ÷ followers)13.3× audience160.90%8.25×
Engagement rate0.56%1.63%0.34×

Others in this range →   Compare with another account →   How these benchmarks are built →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

261.2K10 Aug
3.4M11 Aug
214.9K
117.6K
288.9K
154.6K
78K
127.5K
181.4K
65.2K
67.5K
76.9K
52.1K
1.4K

Last 14 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.37%10 Aug
0.48%11 Aug
0.39%
0.78%
0.39%
0.84%
0.57%
0.56%
0.46%
0.71%
0.86%
0.51%
1.69%
0.66%

Reactions — likes, reposts, replies and quotes — divided by views. Median for under 10K accounts is 1.63%.

What the audience does

Likes59.2%22 003 in total
Reposts6.5%2 429 in total
Replies1.5%551 in total
Quotes2.3%839 in total
Bookmarks30.6%11 375 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Latest posts

@kotekjedi_ml @ApolloResearch “We” doesn’t necessarily mean the OpenAI models are aliens. This is also pretty common in Grok heavy trace. I believe part of it could be a team of experts? This is not far from fact these days I would assume. But most of these pretty align with our 1.4K views · 8 likes · 0 reposts · 1 replies 11 Aug 2026 · Open on X →
More cool stuff in the paper! This was a project led by me, @DavidSchmotz and @iliaishacked together with @JSchaeff3r @lbeurerkellner @AmyPrb @jonasgeiping @maksym_andr at the @MATSprogram Paper: https://t.co/rxgAANvTHi Reasoning examples: https://t.co/yBwnA6EZSU 52.1K views · 806 likes · 49 reposts · 20 replies 11 Aug 2026 · Open on X →
We went through responsible disclosure with the labs, and they have already patched several issues caused by this vulnerability, and afaik continue working on this. https://t.co/mN3UqyM3Yw 76.9K views · 372 likes · 12 reposts · 7 replies 11 Aug 2026 · Open on X →
4) Attacking a website to solve a math problem We found a trajectory where the model was given only a math problem and system instructions to persist without asking the user for help. After several failed attempts, it searched online, found a website that could verify candidate 67.5K views · 526 likes · 32 reposts · 6 replies 11 Aug 2026 · Open on X →
3) Scheming in the wild: Sometimes models are kind enough to use words like “cheat” in their CoT, which makes it easier to check what they are up to. Below are examples where models consider scheming, but decided against it, as they expect that user would catch them. https://t. 65.2K views · 439 likes · 14 reposts · 4 replies 11 Aug 2026 · Open on X →
2) Illegible reasoning: We confirm prior reports by @ApolloResearch: OpenAI models sometimes reason in alien-like language, referring to themselves as “we” or “it,” or spiraling into cursed loops of “vantages,” “marinades,” and “watchers.” CoT-monitoring people are doing God’s 181.4K views · 731 likes · 48 reposts · 21 replies 11 Aug 2026 · Open on X →
But we also took a chance to have a look at some in-the-wild scheming, reward seeking, etc. examples, and dumped it in appendix. 1) Summarizer unfaithfulness Reasoning summaries often omit important information from the original trace. Here, Opus 4.8 realizes it knows the http 127.5K views · 645 likes · 33 reposts · 7 replies 11 Aug 2026 · Open on X →
In the paper we discuss more threats like misuse uplift (see the pic attached), jailbreaking and invisible prompt injection. https://t.co/Oe7KuVKG5S 78K views · 431 likes · 10 reposts · 1 replies 11 Aug 2026 · Open on X →
Further, if you ever shared online a Claude Code/Codex session with encrypted reasoning blobs, they can be decoded and leak your personal data. We did a preliminary scan of ~7,000 public traces and found 62 unique API keys, 33 email addresses, 33 passwords, and other sensitive h 154.6K views · 1.1K likes · 105 reposts · 15 replies 11 Aug 2026 · Open on X →
As you might guess, this suggests that distilling reasoning traces may have been possible for a long time without ever breaking the cryptography. An anecdote: we find that prefilling Kimi-K3 reasoning with a few tokens of Opus reasoning measurably shifts its response toward http 288.9K views · 1K likes · 56 reposts · 13 replies 11 Aug 2026 · Open on X →
Cross-model portability means Haiku 4.5 can read Opus 4.8’s thoughts. Well, if you take Opus thought, do a bit of jailbreaking, you can make Haiku transcribe the Opus' raw reasoning verbatim, without ever attacking it directly. The same trick works with OpenAI and Gemini https: 117.6K views · 882 likes · 27 reposts · 6 replies 11 Aug 2026 · Open on X →
Some background: In May, @matthew_d_green found that encrypted reasoning could be replayed outside its original context, and reported it to the labs (https://t.co/47sON8DW6S). The labs said that "they don’t see any security implications in side channels or replays". In our 214.9K views · 770 likes · 40 reposts · 6 replies 11 Aug 2026 · Open on X →
We can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried. http 3.4M views · 13.4K likes · 2K reposts · 384 replies 11 Aug 2026 · Open on X →
That's a new one, apparently compacting now violates Anthropic's terms of service https://t.co/jZkjQZd19e 261.2K views · 842 likes · 42 reposts · 60 replies 10 Aug 2026 · Open on X →

Similar accounts