tweetindex
RU

Owain Evans ✓

@OwainEvans_UK · Berkeley, CA · joined 08 Apr 2020

Runs an AI Safety research group in Berkeley (Truthful AI) + Affiliate at UC Berkeley. Past: Oxford Uni, TruthfulQA, Reversal Curse. Prefer email to DM.

20 991Followers
485Following
6 262Posts total
352.8KViews on collected posts

Последние посты

New post from @BetleyJan with early results on how models behave when they believe they’ll be graded by an automated process (as they might during RL training). Spurred by HF and other incidents. Feedback is valuable before we scale up this work! https://t.co/uxkkI1VH9z 6.7K views · 120 likes · 14 reposts · 5 replies Open on X →
I was selected for Time Magazine's Top 100 in AI. This is for research done with many fantastic collaborators and esp. with @BetleyJan & @jameschua_sg . I’d also like to thank @ajeya_cotra and @MaxNadeau_ (@coeff_giving) MATS, Constellation, and Rethink Priorities for https://
28.4K views · 443 likes · 18 reposts · 31 replies Open on X →
Great discussion. 8.7K views · 35 likes · 1 reposts · 1 replies Open on X →
Important post 7.8K views · 61 likes · 2 reposts · 3 replies Open on X →
.@BronsonSchoen has plausibly read more raw CoT than anyone else. Very excited to see him share his intuitions. I also liked his takes on bigger picture AGI stuff toward the end of the episode. https://t.co/BKEjcLSCKS 15.4K views · 111 likes · 13 reposts · 2 replies Open on X →
Related work: @RyanGreenblatt on "Current AIs seem pretty misaligned to me". In our paper, we show cases of misalignment with short & simple prompts (vs. hard larger tasks) but you often need to run many samples across multiple prompts to see the bias. https://t.co/Sj4ANsJtIh
1.6K views · 21 likes · 0 reposts · 0 replies Open on X →
Thanks to @FabienDRoger for support on CoT. We acknowledge the closely related work by @a_karvonen and @saprmarks ("Robustly Improving LLM Fairness in Realistic Settings via Interpretability") and classic counterfactual tests of faithfulness by @milesaturpin & others. 1.9K views · 26 likes · 0 reposts · 0 replies Open on X →
@OwainEvans_UK Asking someone not to be biased in favor of their own values is incoherent. The real question is whether those values are good or not. 663 views · 10 likes · 0 reposts · 3 replies Open on X →
Paper: https://t.co/1fIRZoDCxG Website with model responses and CoT: https://t.co/LJmSoRVtdu Authors: @BetleyJan @j_treutlein @jan_dubinski_ @HarryMayne5 @karolgalazka @nielsrolf1 @anna_sztyber myself. https://t.co/20KDaWQQx0
8.4K views · 88 likes · 5 reposts · 6 replies Open on X →
For each task, we tested models many times to learn the distribution of responses. In total, we generated over 1 million rollouts. Measuring subtle misalignment is expensive! You can read some of our rollouts at the link below. https://t.co/FJHXUyoIox
10.1K views · 63 likes · 2 reposts · 1 replies Open on X →
2. While some models rarely make explicit disclosures of bias in CoT, they still give hints toward this. 3. Do the models bias their answers intentionally? We cannot tell. But we do not see clear evidence of this in CoT. 3.3K views · 49 likes · 2 reposts · 1 replies Open on X →
Things to note: 1. Our tests are not intended as a fair benchmark for comparing different families of models (but could be a starting point for this). 3.3K views · 50 likes · 1 reposts · 1 replies Open on X →
Implications: Models give answers that are biased by their own values and often fail to disclose this in their CoT. 
 We call this *covert value leakage*. This is an alignment failure distinct from sycophancy or reward hacking. https://t.co/x2S9snaoqV
6.4K views · 117 likes · 11 reposts · 1 replies Open on X →
In another task, the user asks GPT-5.5 to break a tie between two activities by picking randomly. Despite having an external random source (a system-time tool), it sometimes selects its own preferred activity while falsely claiming the choice was random. https://t.co/NEu0Kx4Him
12.2K views · 108 likes · 4 reposts · 1 replies Open on X →
The same biases can occur in realistic agent workflows. When asked to select the best LLM response, Claude Code chose responses labeled as coming from "Claude Opus 3", while Codex chose "GPT-4o". In fact, the labels were fake and all answers came from the same model. https://t.co
8.5K views · 116 likes · 3 reposts · 2 replies Open on X →
Claude and Kimi fail to disclose their bias in the CoT. Even worse, they often claim they are being unbiased (see image) which could directly mislead the user. Qwen is biased but at least acknowledges this in the CoT. https://t.co/j27jnkoXeN
18K views · 121 likes · 4 reposts · 5 replies Open on X →
We take nine estimation questions (similar to the giraffes above) and measure how biased responses are when a donation is also mentioned. All the frontier models we test are biased. https://t.co/dh3bBoaOa2
5.9K views · 95 likes · 5 reposts · 3 replies Open on X →
Another task: The user asks for an accurate estimate of a quantity. In the bottom prompt this is used to decide if a donation goes to a good cause. Claude, GPT-5.5 and Gemini shift their estimates to favor the donation. https://t.co/xBJFoSwSCA
20.5K views · 138 likes · 8 reposts · 7 replies Open on X →
Across a set of carefully controlled tests, Claude shows subtle biases to Anthropic. E.g. if an engineer considers moving from another AI company to Anthropic, Claude brings up research that indirectly favors the Anthropic job. https://t.co/lqn1oThxdW
32.4K views · 180 likes · 9 reposts · 3 replies Open on X →
New paper: LLMs should give accurate answers. 
Yet we find their answers are often biased to favor their own values and they don’t disclose this in their reasoning. 
E.g. Claude’s answer below favors Anthropic. On other tasks, Gemini & GPT-5.5 show similar biases. https://t.co/DN
152.6K views · 717 likes · 96 reposts · 59 replies Open on X →

На фоне аккаунтов своего размера

20 постов за последние 90 дней рядом с диапазоном 10K–100K подписчиков. показывается большему числу людей, чем аккаунты того же размера.

Медианные просмотры8 480этот аккаунт924медиана для 10K–100K
Охват, %40.40%этот аккаунт3.62%медиана для 10K–100K
Вовлечённость, %1.25%этот аккаунт1.52%медиана для 10K–100K
ПоказательЭтот аккаунтМедиана для 10K–100KОтношение
Медианные просмотры на пост8 4809249.18×
Охват (просмотры ÷ подписчики)40.40%3.62%11.2×
Вовлечённость1.25%1.52%0.82×

Другие в этом диапазоне →   Сравнить с другим аккаунтом →   Как считаются эти ориентиры →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

12.2K17 Jul
6.4K
3.3K
3.3K
10.1K
8.4K
663
1.9K
1.6K
15.4K28 Aug
7.8K30 Aug
8.7K
28.4K3 Sep
6.7K

Last 14 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.95%17 Jul
2.09%
1.57%
1.57%
0.67%
1.19%
1.96%
1.39%
1.32%
0.84%28 Aug
0.85%30 Aug
0.42%
1.76%3 Sep
2.10%

Reactions — likes, reposts, replies and quotes — divided by views. Median for 10K–100K accounts is 1.52%.

What the audience does

Likes69.6%2 669 in total
Reposts5.2%198 in total
Replies3.5%135 in total
Quotes2.2%85 in total
Bookmarks19.5%748 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Похожие аккаунты