tweetindex

Ge Yan

@GeYan_21

CS PhD Student @UW | Previously @UCSanDiego | Intern @ToyotaResearch @allen_ai

856Followers
697Following
141Posts total
128.5KViews on collected posts

Against accounts of the same size

10 posts from the last 90 days, next to the under 10K follower range. ordinary reach for its size, weaker reaction than most.

Median views1 239this account2 112median for under 10K
Reach, %144.74%this account157.00%median for under 10K
Engagement, %0.69%this account1.62%median for under 10K
MetricThis accountMedian for under 10KRatio
Median views per post1 2392 1120.59×
Reach (views ÷ followers)144.74%157.00%0.92×
Engagement rate0.69%1.62%0.43×

Others in this range →   Compare with another account →   How these benchmarks are built →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

113.9K12 Aug
4.3K
2.2K
1.7K
1.3K
1.1K
791
1K
1K
1.1K

Last 10 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.61%12 Aug
0.87%
0.63%
0.71%
0.60%
0.78%
0.63%
0.69%
0.69%
1.61%

Reactions — likes, reposts, replies and quotes — divided by views. Median for under 10K accounts is 1.62%.

What the audience does

Likes51.7%703 in total
Reposts5.3%72 in total
Replies1.2%17 in total
Quotes1.2%16 in total
Bookmarks40.6%553 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Latest posts

That's all! Flex-π matches or beats the strongest VLA in fast action-only mode, and the strongest WAM at full joint generation — one checkpoint, across RoboTwin, LIBERO and a real bimanual YAM robot. We'll be releasing everything: the full codebase, our datasets, and several 1.1K views · 17 likes · 0 reposts · 0 replies 12 Aug 2026 · Open on X →
One more thing: - The result we find most interesting: the extra streams pay off even in the cheapest inference mode. Observing video → +DINO → +pointmap moves RoboTwin success 40% → 47% → 67%, and removing cross-modality forcing costs 21%. - You can also drop the depth 1K views · 6 likes · 0 reposts · 1 replies 12 Aug 2026 · Open on X →
Data Efficiency: All of that from very little data. Fine-tuned on half the real-world demos, Flex-π still beats every baseline trained on all of them. RoboTwin at 50 demos/task: 78.8%, against 41.9% for Fast-WAM, 31.4% for π₀.₅ and 17.2% for LingBot-VA. That's 1.9× the https:// 1K views · 6 likes · 0 reposts · 1 replies 12 Aug 2026 · Open on X →
@GeYan_21 Like like you seems to consider "predicting next action without predicting next frame" as VLA. This is not. By definitition VLA is derived from LLM. What's you'r doing what ever is the regime is a WAM 791 views · 4 likes · 0 reposts · 1 replies 12 Aug 2026 · Open on X →
Generalization: we test all models under significant distribution shift, including novel distractors filling the workspace, and object types the policy has never handled. Flex-π holds 95.0% task completion on Put Plate and 70.0% on Sort Utensils — 2.5 and 5.0 points below its ht 1.1K views · 8 likes · 0 reposts · 1 replies 12 Aug 2026 · Open on X →
Dexterity: a deformable pouch. Unzip it, hold the mouth open, drop a pen in, re-grasp the pull and close it — six steps, each reachable only if the last worked, on an object with no stable shape. Task completion: Flex-π 70.0%, π₀.₅ 42.8%, ManiFlow 31.9%. https://t.co/OnEDGhd5J0 1.3K views · 7 likes · 0 reposts · 1 replies 12 Aug 2026 · Open on X →
On 5 dexterous bimanual YAM tasks, Flex-π beats π₀.₅, Fast-WAM and ManiFlow on every single one — 2.3× the success rate of the strongest baseline. Even action-only, the cheapest policy here to run, is ahead of all of them. Precision: the robot repairs its own gripper. Eight http 1.7K views · 11 likes · 0 reposts · 1 replies 12 Aug 2026 · Open on X →
Then we randomly drop visual streams during training, and force the model to generate the ones it never saw as input. We call this cross-modality forcing: imagine future 3D geometry with no pointmap given, future semantics from RGB and geometry alone, and so on. To do that, the h 2.2K views · 13 likes · 0 reposts · 1 replies 12 Aug 2026 · Open on X →
Everyone's training WAMs now, and most of them predict one thing: future RGB latents. Video is a strong prior, but it's trained to reconstruct pixels, and carries no explicit signal for the accurate 3D geometry or object semantics that manipulation actually needs. Adding those h 4.3K views · 33 likes · 1 reposts · 2 replies 12 Aug 2026 · Open on X →
Are VLAs dead? No, and you don't have to choose between VLAs and WAMs. Introducing Flex-π: a multi-stream world-action model (WAM) that jointly predicts future RGB, 3D pointmaps and DINO semantics with actions in training, then deploys as a VLA, a full WAM, or anything in https 113.9K views · 598 likes · 71 reposts · 8 replies 12 Aug 2026 · Open on X →

Similar accounts