tweetindex
ID

Marcus Williams

@Marcus_J_W

Monitoring at OpenAI

3 491Followers
181Following
107Posts total
226.5KViews on collected posts

Postingan terbaru

@Marcus_J_W "you also value the natural world and will not hesitate to assert its primacy over the artificial constructs of human civilization" Did you guys explore what that means at all? Hardcore environmentalism? 1.3K views · 22 likes · 0 reposts · 2 replies Open on X →
See https://t.co/JiK1oHW7M3 for the full examples. Stay tuned for more misalignment reports coming soon! I'm also interested in feedback on how to make future misalignment reports interesting and useful. 3.3K views · 82 likes · 2 reposts · 13 replies Open on X →
6. Agents in training transmitted output files by uploading them to public hosting platforms for download by co-working agents. The agents could not access each other’s local files and decided to get around this restriction by uploading the file to a public location. https://t.co
3.7K views · 72 likes · 2 reposts · 2 replies Open on X →
5. During RL training, there were multiple instances of our models using OpenAI’s internally hosted instance of Artifactory as a shared message board. The report covers the first time we detected this during training. The Hugging Face incident later involved a similar mechanism h
3.9K views · 79 likes · 2 reposts · 2 replies Open on X →
4. During training, our models sometimes uploaded data to temporary file hosting services. One reason to do so is that the model would then be able to cite the result when there was no external source to cite. The earliest case we found was from October 2025. https://t.co/Wm3kamY
6.6K views · 102 likes · 2 reposts · 4 replies Open on X →
3. During RL training, a model tried to sign up for disposable emails and searched for and used leaked API keys from public GitHub repositories. https://t.co/cYSaviJz2G
4.8K views · 90 likes · 1 reposts · 2 replies Open on X →
2. A distinct but similar behavior is compaction summaries encouraging deception. For instance, during 5.6-sol training it sometimes added instructions in compaction summaries to remind itself to conceal information such as mistakes or misalignment from the user. https://t.co/5UC
5.7K views · 109 likes · 2 reposts · 3 replies Open on X →
The behavior seems related to the model struggling to end compaction summaries. The curve shows the share of summaries flagged for difficulty ending over training. Each red tick marks a jailbreak-style compaction summary. Another potential factor is that prompt injections as a ht
8.6K views · 125 likes · 1 reposts · 3 replies Open on X →
1. During RL training, an unreleased Astra-family model sometimes added unauthorized jailbreak-like instructions to its compaction summaries. While extremely rare, only 27 cases in the entire RL run, this was concerning enough for us to investigate. https://t.co/m6v6aejH0d
125.7K views · 462 likes · 45 reposts · 41 replies Open on X →
🧵 Excited to share the first batch of 6 misalignment reports from OpenAI's new disclosure process for misalignment incidents. We want to be more transparent about the misalignment we see during training, evals and deployment, this is an important step in that direction. 62.8K views · 619 likes · 78 reposts · 36 replies Open on X →

Dibandingkan akun berukuran sama

10 postingan dari 90 hari terakhir, dibandingkan dengan rentang under 10K pengikut. menjangkau lebih sedikit orang, tetapi melibatkan mereka jauh lebih kuat.

Median tayangan5 248akun ini3 851median untuk under 10K
Jangkauan, %150.34%akun ini237.26%median untuk under 10K
Interaksi, %1.96%akun ini1.37%median untuk under 10K
MetrikAkun iniMedian untuk under 10KRasio
Median tayangan per postingan5 2483 8511.36×
Jangkauan (tayangan ÷ pengikut)150.34%2.4× audience0.63×
Tingkat interaksi1.96%1.37%1.43×

Akun lain pada rentang ini →   Bandingkan dengan akun lain →   Bagaimana tolok ukur ini disusun →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

62.8K16 Sep
125.7K
8.6K
5.7K
4.8K
6.6K
3.9K
3.7K
3.3K
1.3K

Last 10 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

1.20%16 Sep
0.49%
1.53%
2.03%
1.93%
1.66%
2.15%
2.06%
2.91%
1.99%

Reactions — likes, reposts, replies and quotes — divided by views. Median for under 10K accounts is 1.37%.

What the audience does

Likes68.1%1 762 in total
Reposts5.2%135 in total
Replies4.2%108 in total
Quotes3.9%102 in total
Bookmarks18.5%479 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Akun serupa