tweetindex
RU

Tony Wu ✓

@tonywu_71

Multimodal, RAG, Agents | ColPali co-first author | @centralesupelec 🇫🇷 x @Cambridge_Uni 🇬🇧 | Core Researcher at @hcompany_ai 🧑🏻‍💻

1 565Followers
544Following
572Posts total
29.3KViews on collected posts

Последние посты

@tonywu_71 congrats Tony! very interesting work and the performance seems to be strong! And i like a lot you put a lot of efforts on the writing and making beautiful charts in the report! 242 views · 9 likes · 0 reposts · 1 replies Open on X →
And many thanks to Raushan Turganbay, @mervenoyann and @art_zucker for helping me land the day-0 transformers integration! And huge kudos to @tomaarsen, the goat maintainer of Sentence Transformers who recently added late-interaction support! 🤗 (15/N) 232 views · 9 likes · 0 reposts · 0 replies Open on X →
Cc to the late-interaction gang 🙏🏼 @lateinteraction @ManuelFaysse @bclavie @antoine_chaffin @raphaelsrty @AmelieTabatta @MaceQuent1 @pteiletche @antoine_edy @bo_wangbo @jobergum @doesdatmaksense @helloiamleonie @CShorten30 (14/N) 376 views · 13 likes · 0 reposts · 1 replies Open on X →
NeoMME began as a side quest between two good friends (thanks @Aurelien_L_ for bearing with me 🫶🏼). We both worked with limited time, but I hope you'll find our work interesting! And thanks to @hcompany_ai for supporting this work and providing the training compute. (13/N) 244 views · 12 likes · 0 reposts · 1 replies Open on X →
We wrote a longer blog post and a 54-page tech report to share all our learnings with the community. You can find these in our 🤗 Hf Collection, which includes the Hf Space demo and the open-weight Apache 2.0 checkpoints. (12/N) https://t.co/1968WUaj3l 274 views · 17 likes · 0 reposts · 1 replies Open on X →
You can use NeoMME-Retriever for visual RAG: given a query, it retrieves the top-k most relevant page images. A VLM then answers the original question using those pages as context. I created an Hf Space so you can try it yourself! (11/N) https://t.co/8pARXLHT1s
0:19
271 views · 13 likes · 0 reposts · 1 replies Open on X →
NeoMME-Retriever is available in 🤗 Transformers. The snippet below scores two queries against two dummy page images using both dense and late-interaction heads. (10/N) https://t.co/x2dKbZjM8X
286 views · 15 likes · 0 reposts · 1 replies Open on X →
For NeoMME-Retriever-260M on ViDoRe v3, high-resolution late-interaction embeddings average about 1.5 MB per page. Token pooling and asymmetric quantization reduce them to 6 kB, 255x smaller, while retaining more than 95% of baseline nDCG@10. (9/N) https://t.co/6c1Z164NoL
1.2K views · 18 likes · 1 reposts · 1 replies Open on X →
With matched 2048x2048 inputs on one NVIDIA L40S, NeoMME-Retriever-260M encodes about 51 pages per second, nearly 2x ColModernVBERT's 26 pages per second. Faster document encoding means less GPU time, so cheaper indexing! ⚡️ (8/N) https://t.co/F6zOncFPiK
336 views · 15 likes · 0 reposts · 1 replies Open on X →
On ViDoRe v3 and using the late-interaction embeddings, NeoMME-Retriever-260M reaches 0.523 nDCG@10, the highest score among evaluated models below 800M parameters. The 800M model reaches 0.556. Both are on the benchmark's model-size Pareto frontier. (7/N) https://t.co/RnpQVmJwII
1.3K views · 16 likes · 1 reposts · 1 replies Open on X →
To get a meaningful downstream evaluation of NeoMME, we fine-tuned the model for visual document retrieval 🔎 We added two jointly trained heads so the model can output both dense and late-interaction embeddings in a single forward pass. (6/N) https://t.co/p3ZAiwPogj
389 views · 12 likes · 0 reposts · 1 replies Open on X →
We pretrain NeoMME from scratch with masked discrete diffusion. Image patches stay visible while the model reconstructs masked text. At high masking rates, the model is forced to use the image instead of guessing from the visible text. (5/N) https://t.co/jIuHrvbTmy
433 views · 15 likes · 0 reposts · 2 replies Open on X →
NeoMME splits images into 32x32 patches and keeps their aspect ratio and resolution. Most layers use sliding-window attention, while every sixth layer and the final layer use global attention. A 16k-token context fits up to two 4K UHD images in one forward pass. (4/N) https://t.c
485 views · 15 likes · 0 reposts · 1 replies Open on X →
With NeoMME, we decided to build a single bidirectional Transformer encoder that processes both text tokens and raw image patches and could be trained entirely from scratch. (3/N) https://t.co/DN7H5UWSZO
776 views · 17 likes · 0 reposts · 1 replies Open on X →
Multimodal encoders like ColPali generally reuse pretrained generative VLMs architecture (vision encoder → projector → LLM). Such models are powerful, but they are usually large and the asymmetry in how modalities are handled make it suboptimal for deployment, fine-tuning, 927 views · 18 likes · 0 reposts · 1 replies Open on X →
👀 Meet NeoMME: a family of 260M and 800M Multimodal-Native Multilingual efficient Encoders One bidirectional Transformer processes text tokens and raw image patches, with no pretrained vision tower, text encoder, or decoder. (1/N 🧵) https://t.co/0SFjprphH8 21.6K views · 168 likes · 30 reposts · 11 replies Open on X →

На фоне аккаунтов своего размера

16 постов за последние 90 дней рядом с диапазоном under 10K подписчиков. доходит до меньшего числа людей, зато вовлекает их куда сильнее.

Медианные просмотры382этот аккаунт4 175медиана для under 10K
Охват, %24.44%этот аккаунт250.82%медиана для under 10K
Вовлечённость, %3.80%этот аккаунт1.41%медиана для under 10K
ПоказательЭтот аккаунтМедиана для under 10KОтношение
Медианные просмотры на пост3824 1750.09×
Охват (просмотры ÷ подписчики)24.44%2.5× audience0.10×
Вовлечённость3.80%1.41%2.70×

Другие в этом диапазоне →   Сравнить с другим аккаунтом →   Как считаются эти ориентиры →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

7763 Sep
485
433
389
1.3K
336
1.2K
286
271
274
244
376
232
242

Last 14 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

2.32%3 Sep
3.30%
3.93%
3.34%
1.42%
4.76%
1.69%
5.59%
5.17%
6.57%
5.33%
3.72%
3.88%
4.13%

Reactions — likes, reposts, replies and quotes — divided by views. Median for under 10K accounts is 1.41%.

What the audience does

Likes68.5%382 in total
Reposts5.7%32 in total
Replies4.7%26 in total
Bookmarks21.1%118 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Похожие аккаунты