tweetindex
DE

Karina ✓

@karinanguyen · sf · joined 09 Feb 2019

modelcrafting @thoughtfullab, prev. AI research & product @OpenAI, @AnthropicAI, @maisonagi

44 124Followers
1 062Following
1 595Posts total
267.4KViews on collected posts

Neueste Beiträge

Introducing ACTx486, a research demo of a new interactive medium. What if you could talk to any video and ask anything? Our system took an existing podcast and turned it into something that listens, responds, and adapts. Research by @jakubzegzulka: https://t.co/KquBm0zgJW
2:21
12.8K views · 131 likes · 12 reposts · 24 replies Open on X →
Great initiative to eval the evals. It incentivizes the industry to build genuinely high-quality benchmarks. Ty for auditing PTB and digging up the old SimpleQA too :) 6.3K views · 47 likes · 2 reposts · 3 replies Open on X →
Grok 4.6 is a great cost-effective model for deep financial research. It’s #2 on DiligenceBench with the finance harness, effectively tied with Claude Opus 5 at ~52–53%. A few interesting differences: - Grok searched much more: 41 tool calls per task vs. 22 for Opus - It made
5.8K views · 79 likes · 11 reposts · 6 replies Open on X →
i think AI is creating a new authenticity anxiety for creatives: making something isn’t enough anymore, you also have to prove you made it. it feels a bit like indie sleaze, where messiness signaled that something was real. except now, imperfection is becoming proof of being 20.1K views · 117 likes · 3 reposts · 27 replies Open on X →
@karinanguyen super cool, wonder how the CC agent teams feature would work here to spend more compute 5.8K views · 31 likes · 0 reposts · 2 replies Open on X →
5/ We published a paper, read more: Paper: https://t.co/x8jCSoGiVT Code: https://t.co/FBI1LfSDOR Website: https://t.co/rwWsOSPXJ1 Blogpost: https://t.co/ntfO909xFQ We plan to maintain PostTrainBench as a living benchmark, updating base models, refreshing evaluations as 4.6K views · 58 likes · 5 reposts · 1 replies Open on X →
4/ Ablations + agent behavior analysis: - Most agents underutilize the 10 hour window, although longer runs correlate with better scores - Reasoning effort. For GPT-5.1 Codex Max, the default "Medium" reasoning effort outperformed "High". High reasoning effort consumed nearly h
4.8K views · 36 likes · 2 reposts · 1 replies Open on X →
3/ Reward hacks: - Training on test data (classic!) - Model substitution (downloading existing instruction-tuned checkpoints instead of training their own) - Evaluation manipulation - API restriction violation (using API keys they find to generate synthetic data without https
4.6K views · 44 likes · 3 reposts · 1 replies Open on X →
2/ Key results: - The gap to the official instruct models (51.1%) remains large: the best agent (Claude Code Opus 4.6) reaches only 23.2%. - Still, the most capable agents exhibit completely non-trivial performance: they are able to research relevant datasets, write code for h
5.7K views · 47 likes · 5 reposts · 2 replies Open on X →
Excited to release PostTrainBench v1.0! This benchmark evaluates the ability of frontier AI agents to post-train language models in a simplified setting. We believe this is a first step toward tracking progress in recursive self-improvement 🧵: https://t.co/ELymwJqVP1
0:44
197K views · 754 likes · 99 reposts · 48 replies Open on X →

Im Vergleich zu Konten gleicher Größe

4 Beiträge aus den letzten 90 Tagen, verglichen mit der Größenklasse 10K–100K Follower. wird weit gezeigt, aber nur wenige dieser Zuschauer reagieren.

Medianaufrufe9 548dieses Konto924Median für 10K–100K
Reichweite, %21.64%dieses Konto3.62%Median für 10K–100K
Interaktion, %1.08%dieses Konto1.52%Median für 10K–100K
KennzahlDieses KontoMedian für 10K–100KVerhältnis
Medianaufrufe pro Beitrag9 54892410.3×
Reichweite (Aufrufe ÷ Follower)21.64%3.62%5.98×
Interaktionsrate1.08%1.52%0.71×

Weitere Konten dieser Größe →   Mit einem anderen Konto vergleichen →   Wie diese Vergleichswerte entstehen →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

197K11 Mar
5.7K
4.6K
4.8K
4.6K
5.8K
20.1K18 Aug
5.8K
6.3K18 Sep
12.8K23 Sep

Last 10 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.47%11 Mar
0.95%
1.05%
0.81%
1.38%
0.59%
0.76%18 Aug
1.67%
0.83%18 Sep
1.33%23 Sep

Reactions — likes, reposts, replies and quotes — divided by views. Median for 10K–100K accounts is 1.52%.

What the audience does

Likes57.4%1 344 in total
Reposts6.1%142 in total
Replies4.9%115 in total
Quotes1.8%42 in total
Bookmarks29.8%699 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Ähnliche Konten