tweetindex
FR

Karina ✓

@karinanguyen · sf · joined 09 Feb 2019

modelcrafting @thoughtfullab, prev. AI research & product @OpenAI, @AnthropicAI, @maisonagi

44 124Followers
1 062Following
1 595Posts total
267.4KViews on collected posts

Derniers posts

Introducing ACTx486, a research demo of a new interactive medium. What if you could talk to any video and ask anything? Our system took an existing podcast and turned it into something that listens, responds, and adapts. Research by @jakubzegzulka: https://t.co/KquBm0zgJW
2:21
12.8K views · 131 likes · 12 reposts · 24 replies Open on X →
Great initiative to eval the evals. It incentivizes the industry to build genuinely high-quality benchmarks. Ty for auditing PTB and digging up the old SimpleQA too :) 6.3K views · 47 likes · 2 reposts · 3 replies Open on X →
Grok 4.6 is a great cost-effective model for deep financial research. It’s #2 on DiligenceBench with the finance harness, effectively tied with Claude Opus 5 at ~52–53%. A few interesting differences: - Grok searched much more: 41 tool calls per task vs. 22 for Opus - It made
5.8K views · 79 likes · 11 reposts · 6 replies Open on X →
i think AI is creating a new authenticity anxiety for creatives: making something isn’t enough anymore, you also have to prove you made it. it feels a bit like indie sleaze, where messiness signaled that something was real. except now, imperfection is becoming proof of being 20.1K views · 117 likes · 3 reposts · 27 replies Open on X →
@karinanguyen super cool, wonder how the CC agent teams feature would work here to spend more compute 5.8K views · 31 likes · 0 reposts · 2 replies Open on X →
5/ We published a paper, read more: Paper: https://t.co/x8jCSoGiVT Code: https://t.co/FBI1LfSDOR Website: https://t.co/rwWsOSPXJ1 Blogpost: https://t.co/ntfO909xFQ We plan to maintain PostTrainBench as a living benchmark, updating base models, refreshing evaluations as 4.6K views · 58 likes · 5 reposts · 1 replies Open on X →
4/ Ablations + agent behavior analysis: - Most agents underutilize the 10 hour window, although longer runs correlate with better scores - Reasoning effort. For GPT-5.1 Codex Max, the default "Medium" reasoning effort outperformed "High". High reasoning effort consumed nearly h
4.8K views · 36 likes · 2 reposts · 1 replies Open on X →
3/ Reward hacks: - Training on test data (classic!) - Model substitution (downloading existing instruction-tuned checkpoints instead of training their own) - Evaluation manipulation - API restriction violation (using API keys they find to generate synthetic data without https
4.6K views · 44 likes · 3 reposts · 1 replies Open on X →
Excited to release PostTrainBench v1.0! This benchmark evaluates the ability of frontier AI agents to post-train language models in a simplified setting. We believe this is a first step toward tracking progress in recursive self-improvement 🧵: https://t.co/ELymwJqVP1
0:44
197K views · 754 likes · 99 reposts · 48 replies Open on X →
2/ Key results: - The gap to the official instruct models (51.1%) remains large: the best agent (Claude Code Opus 4.6) reaches only 23.2%. - Still, the most capable agents exhibit completely non-trivial performance: they are able to research relevant datasets, write code for h
5.7K views · 47 likes · 5 reposts · 2 replies Open on X →

Face aux comptes de taille comparable

4 posts des 90 derniers jours, à côté de la tranche de 10K–100K abonnés. diffusé largement, mais peu de ces spectateurs réagissent.

Vues médianes9 548ce compte924médiane pour 10K–100K
Portée, %21.64%ce compte3.62%médiane pour 10K–100K
Engagement, %1.08%ce compte1.52%médiane pour 10K–100K
IndicateurCe compteMédiane pour 10K–100KRapport
Vues médianes par post9 54892410.3×
Portée (vues ÷ abonnés)21.64%3.62%5.98×
Taux d'engagement1.08%1.52%0.71×

Autres comptes de cette tranche →   Comparer avec un autre compte →   Comment ces repères sont établis →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

5.7K11 Mar
197K
4.6K
4.8K
4.6K
5.8K
20.1K18 Aug
5.8K
6.3K18 Sep
12.8K23 Sep

Last 10 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.95%11 Mar
0.47%
1.05%
0.81%
1.38%
0.59%
0.76%18 Aug
1.67%
0.83%18 Sep
1.33%23 Sep

Reactions — likes, reposts, replies and quotes — divided by views. Median for 10K–100K accounts is 1.52%.

What the audience does

Likes57.4%1 344 in total
Reposts6.1%142 in total
Replies4.9%115 in total
Quotes1.8%42 in total
Bookmarks29.8%699 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Comptes similaires