tweetindex
EN

Thomas G. Dietterich

@tdietterich · Corvallis, OR · joined 19 Aug 2012

University Distinguished Professor (Emeritus), Oregon State Univ.; Former President, AAAI; ArXiv Editor in Chief & Chair CS Section

64 027Followers
659Following
17 583Posts total
1.1MViews on collected posts

Latest posts

@tdietterich @JessicaHullman Relevant review: https://t.co/emrOqEMpgT 631 views · 2 likes · 0 reposts · 0 replies Open on X →
This is no guarantee (as there is no guarantee of absolute safety), but it is an important and necessary shift in how we think about the future of "autonomy". end/ 1K views · 25 likes · 1 reposts · 0 replies Open on X →
We should test these human-machine systems using our best adversarial methods and measure the time to detect a failure, the time to recover from it, and the cost of the (hopefully, simulated) harms. 13/ 1.1K views · 20 likes · 1 reposts · 1 replies Open on X →
The experience with "autonomous" vehicles and the recent AI cyber security incidents from OpenAI and Google demonstrate that our AI systems are not different. Waymo requires a team of human supervisors. 10/ 1.1K views · 17 likes · 1 reposts · 3 replies Open on X →
The conclusion is that we must design our "autonomous" systems to be combined human-machine systems with continual supervision from the very beginning. 12/ 2.3K views · 31 likes · 3 reposts · 1 replies Open on X →
The OpenAI agents lacked sufficient supervision and exhibited unanticipated failure modes. Software professionals deploying coding agents are spending most of their time on supervision. 11/ 1.1K views · 19 likes · 1 reposts · 1 replies Open on X →
But adversaries also operate within a closed space of moves, and we rarely have a proof of completeness (i.e., that they will find a failure case if one exists). 8/ 1.2K views · 16 likes · 2 reposts · 1 replies Open on X →
Similarly, simulations are always limited, and they are impossible or unethical to validate in life-threatening scenarios. 9/ 1.1K views · 16 likes · 0 reposts · 1 replies Open on X →
How can we generate a benchmark of novel failures and environmental changes that we, by definition, cannot anticipate? The best tool we have is adversarial challenge in simulation. 7/ 1.3K views · 22 likes · 1 reposts · 3 replies Open on X →
Nancy Leveson drives this point home in her excellent book, Engineering a Safer World. "Safety" (or in today's parlance, "alignment") is not a property of the automation (the self-driving car, the AI agent). It is a dynamic property that must be maintained through active control. 3.2K views · 48 likes · 6 reposts · 2 replies Open on X →
Leveson focuses on the role of government, regulators, owners, and operators where the threats to safety are primarily budget cuts and staff turnover. But she also discusses changes in the operating environment (including regulations). 4/ 1.8K views · 23 likes · 1 reposts · 1 replies Open on X →
Similarly, @ddwoods2 has devoted much of his career to studying "automation surprises", and he argues that we have never created an autonomous system that can adapt to novelty. It is always the human operators who provide the adaptive capability. 5/ 2.6K views · 32 likes · 2 reposts · 2 replies Open on X →
Those of us designing AI agents are hoping (assuming?) that this case will be different -- that these agents will have enough intelligence to reason about novel failures and environmental changes and respond appropriately. However, testing such systems is extremely difficult. 6/ 1.5K views · 26 likes · 3 reposts · 2 replies Open on X →
A lesson from decades of work on automation is that autonomous systems require continual monitoring and human oversight. System designers never(*) succeed in anticipating all of the ways that a system can fail. 1/ 31.4K views · 407 likes · 73 reposts · 16 replies Open on X →
In particular, failure modes of the system components and changes in the operating environment exhibit unbounded variety. (*)The exceptions are simple systems performing very narrow tasks in unchanging environments. 2/ 2.5K views · 41 likes · 1 reposts · 1 replies Open on X →
Maybe author exams *can* scale? Cool work! 6.8K views · 46 likes · 1 reposts · 5 replies Open on X →
🗣️ Can we evaluate the process of intellectual production rather than just the final text? Can we creatively design approaches to filter unaccountable research and nudge researchers and students toward genuine ownership of their work? Yes. 1/11 🧵 https://t.co/pkD10LUted
18.4K views · 121 likes · 24 reposts · 18 replies Open on X →
@tdietterich @arxiv You’re proposals are that only well-connected academics who are fluent in English and comfortable in-person interview be allowed to publish preprints. You’ve critically lost perspective on the purpose of a preprint server. This will serve to further gatekeep a 7.8K views · 148 likes · 0 reposts · 3 replies Open on X →
Another approach might be to conduct an oral examination of the author to test their understanding of the paper. Could this scale, perhaps by requiring authors to provide a letter from a known authority certifying that the authority has conducted such an exam? 4/ 81.2K views · 346 likes · 9 reposts · 22 replies Open on X →
What do people think? 45K views · 208 likes · 2 reposts · 144 replies Open on X →
An alternative is to create a new kind of venue where AI-written and verified results could be submitted and made available to the research community for examination and analysis. Like aiXiv, there would be no human author, only a "corresponding human". 5/ 57.1K views · 541 likes · 23 reposts · 34 replies Open on X →
ArXiv policy is that a human author must take responsibility for a submission. But obviously an author cannot take responsibility (or claim credit) for a result that they do not understand. 2/ 61.6K views · 455 likes · 16 reposts · 9 replies Open on X →
How should the scientific enterprise handle this? We could try to decide whether the author understands the work. E.g., if the author lacks a publication record or credentials in the research area, we could reject the submission. 3/ 100.7K views · 288 likes · 10 reposts · 37 replies Open on X →
We are seeing a new trend in submissions to @arxiv (and presumably to conferences and journals): Authors submitting papers whose contents they likely do not understand. 1/ 659.2K views · 2.8K likes · 357 reposts · 191 replies Open on X →
Why do we call it Recursive Self Improvement? At most, it is Iterative Self Improvement. RSI makes my wrists hurt... 16K views · 83 likes · 6 reposts · 13 replies Open on X →
I've been trying to imagine how the ML research and publication enterprise could be re-organized. Here are some initial thoughts. Feedback welcome! 1/ 0 views · 459 likes · 71 reposts · 45 replies Open on X →

Against accounts of the same size

25 posts from the last 90 days, next to the 10K–100K follower range. right around the median for its follower range.

Median views2 616this account924median for 10K–100K
Reach, %4.09%this account3.62%median for 10K–100K
Engagement, %1.48%this account1.52%median for 10K–100K
MetricThis accountMedian for 10K–100KRatio
Median views per post2 6169242.83×
Reach (views ÷ followers)4.09%3.62%1.13×
Engagement rate1.48%1.52%0.97×

Others in this range →   Compare with another account →   How these benchmarks are built →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

31.4K20 Sep
1.5K
2.6K
1.8K
3.2K
1.3K
1.1K
1.2K
1.1K
2.3K
1.1K
1.1K
1K
631

Last 14 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

1.63%20 Sep
2.01%
1.38%
1.40%
1.77%
2.01%
1.48%
1.57%
1.96%
1.54%
1.87%
2.06%
2.57%
0.32%

Reactions — likes, reposts, replies and quotes — divided by views. Median for 10K–100K accounts is 1.52%.

What the audience does

Likes69.0%5 761 in total
Reposts6.5%544 in total
Replies6.1%511 in total
Quotes2.6%218 in total
Bookmarks15.8%1 316 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

Followers by day

21 Sep

Daily snapshots since 21 Sep 2026; the dashed line is the starting count.

Similar accounts