tweetindex

Keyon Vafa

@keyonV

Incoming assistant professor @UCBerkeley | Now @Harvard_Data & @bicyclesftmind @MIT | Prev PhD @Columbia | Researching AI/implicit world models

5 436Followers
958Following
1 435Posts total
2.9MViews on collected posts

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

33.4K11 Jul
31.2K
33.1K
90K
35.4K
27.1K
23.1K
25K
21.2K
20.5K
26.4K
18.9K
23.5K
6.3K

Last 14 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.72%11 Jul
0.68%
0.74%
0.37%
0.62%
0.79%
0.92%
0.87%
0.82%
1.44%
0.69%
1.13%
1.37%
0.83%

Reactions — likes, reposts, replies and quotes — divided by views.

What the audience does

Likes65.2%14 597 in total
Reposts6.9%1 554 in total
Replies1.6%348 in total
Quotes2.2%482 in total
Bookmarks24.1%5 397 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Latest posts

@keyonV > Can an AI model predict perfectly and still have a terrible world model? Of course, for a given isolated problem. This is called supervised learning and you can get as perfect as you want to be by providing training examples. 6.3K views · 49 likes · 0 reposts · 3 replies 11 Jul 2025 Paper: https://t.co/DGJvqTWD3O Co-authors: Peter Chang (@petergchang), Ashesh Rambachan (@asheshrambachan), Sendhil Mullainathan (@m_sendhil) 23.5K views · 286 likes · 16 reposts · 19 replies 11 Jul 2025 This is one way to evaluate world models. But there are many other interesting approaches. Plug: If you're interested in more, check out the Workshop on Assessing World Models I'm co-organizing next Friday at ICML. https://t.co/IYKF2RmHn7 18.9K views · 202 likes · 6 reposts · 6 replies 11 Jul 2025 Last year we proposed different tests that studied single tasks. We now think that studying behavior on new tasks better captures what we want from foundation models: tools for new problems. It's what separates Newton's laws from Kepler's predictions. https://t.co/CwqGY6obUV 26.4K views · 176 likes · 4 reposts · 3 replies 11 Jul 2025 Summary: 1. We propose inductive bias probes: a model's inductive bias reveals its world model 2. Foundation models can have great predictions with poor world models 3. One reason world models are poor: models group together distinct states that have similar allowed next-tokens 20.5K views · 274 likes · 11 reposts · 6 replies 11 Jul 2025 Inductive bias probes can test this hypothesis more generally. Models are much likelier to conflate two separate states when they share the same legal next-tokens. https://t.co/sgGFgsYR4O 21.2K views · 168 likes · 2 reposts · 1 replies 11 Jul 2025 We fine-tune an Othello next-token prediction model to reconstruct boards. Even when the model reconstructs boards incorrectly, the reconstructed boards often get the legal next moves right. Models seem to construct "enough of" the board to calculate single next moves. https:// 25K views · 210 likes · 3 reposts · 3 replies 11 Jul 2025 If a foundation model's inductive bias isn't toward a given world model, what is it toward? One hypothesis: models confuse sequences that belong to different states but have the same legal *next* tokens. Example: Two different Othello boards can have the same legal next moves. 23.1K views · 201 likes · 6 reposts · 5 replies 11 Jul 2025 We also apply these probes to lattice problems (think gridworld). Inductive biases are great when the number of states is small. But they deteriorate quickly. Recurrent and state-space models like Mamba consistently have better inductive biases than transformers. https://t.co/b 27.1K views · 206 likes · 5 reposts · 1 replies 11 Jul 2025 Would more general models like LLMs do better? We tried providing o3, Claude Sonnet 4, and Gemini 2.5 Pro with a small number of force magnitudes in-context w/o saying what they are. These LLMs are explicitly trained on Newton's laws. But they can't get the rest of the forces. 35.4K views · 210 likes · 3 reposts · 4 replies 11 Jul 2025 We then fine-tuned the model on a larger scale, to predict forces across 10K solar systems. We used a symbolic regression to compare the recovered force law to Newton's law. It not only recovered a nonsensical law—it recovered different laws for different galaxies. https://t.co 90K views · 298 likes · 19 reposts · 5 replies 11 Jul 2025 To demonstrate, we fine-tuned the model to predict force vectors on a small dataset of planets in our solar system. A model that understands Newtonian mechanics should get these. But the transformer struggles. https://t.co/1sULoGIpHP 33.1K views · 226 likes · 7 reposts · 7 replies 11 Jul 2025 But has the model discovered Newton's laws? When we fine-tune it to new tasks, its inductive bias isn't toward Newtonian states. When it extrapolates, it makes similar predictions for orbits with very different states, and different predictions for orbits with similar states. h 31.2K views · 205 likes · 6 reposts · 2 replies 11 Jul 2025 We apply these probes to orbital, lattice, and Othello problems. Starting with orbits: we encode solar systems as sequences and train a transformer on 10M solar systems (20B tokens) The model makes accurate predictions many timesteps ahead. Predictions for our solar system: htt 33.4K views · 232 likes · 5 reposts · 2 replies 11 Jul 2025 We propose a method to measure these inductive biases. We call it an inductive bias probe. Two steps: 1. Fit a foundation model to many new, very small synthetic datasets 2. Analyze patterns in the functions it learns to find the model's inductive bias https://t.co/AJn2JFMqip 37.5K views · 318 likes · 14 reposts · 1 replies 11 Jul 2025 Newton's laws are a kind of foundation model. They provide a place to start when working on new problems. A good foundation model should do the same. The No Free Lunch Theorem motivates a test: Every foundation model has an inductive bias. This bias reveals its world model. 38.8K views · 313 likes · 8 reposts · 3 replies 11 Jul 2025 If you only care about orbits, Newton didn't add much. His laws give the same predictions. But Newton's laws went beyond orbits: the same laws explain pendula, cannonballs, and rockets. This motivates our framework: Predictions apply to one task. World models generalize to many 41.7K views · 500 likes · 14 reposts · 4 replies 11 Jul 2025 Perhaps the most influential world model had its start as a predictive model. Before we had Newton's laws of gravity, we had Kepler's predictions of planetary orbits. Kepler's predictions led to Newton's laws. So what did Newton add? https://t.co/Wxoc8O9c1E 43.6K views · 334 likes · 12 reposts · 3 replies 11 Jul 2025 Our paper aims to answer two questions: 1. What's the difference between prediction and world models? 2. Are there straightforward metrics that can test this distinction? Our paper is about AI. But it's helpful to go back 400 years to answer these questions. https://t.co/1gCi9M 71.4K views · 567 likes · 30 reposts · 10 replies 11 Jul 2025 Can an AI model predict perfectly and still have a terrible world model? What would that even mean? Our new ICML paper formalizes these questions One result tells the story: A transformer trained on 10M solar systems nails planetary orbits. But it botches gravitational laws 🧵 1.4M views · 6.6K likes · 987 reposts · 210 replies 11 Jul 2025 New paper: How can you tell if a transformer has the right world model? We trained a transformer to predict directions for NYC taxi rides. The model was good. It could find shortest paths between new points But had it built a map of NYC? We reconstructed its map and found this: 841.2K views · 3K likes · 396 reposts · 50 replies 20 Jun 2024

Similar accounts