tweetindex

tomaarsen

@tomaarsen

Sentence Transformers Lead, SetFit & NLTK maintainer Machine Learning Engineer at 🤗 Hugging Face

4 814Followers
479Following
2 331Posts total
141.7KViews on collected posts

Against accounts of the same size

22 posts from the last 90 days, next to the under 10K follower range. reaches fewer people than peers, but engages them much harder.

Median views637this account3 166median for under 10K
Reach, %13.23%this account193.65%median for under 10K
Engagement, %2.61%this account1.50%median for under 10K
MetricThis accountMedian for under 10KRatio
Median views per post6373 1660.20×
Reach (views ÷ followers)13.23%193.65%0.07×
Engagement rate2.61%1.50%1.74×

Others in this range →   Compare with another account →   How these benchmarks are built →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

68618 Aug
731
588
557
544
458
489
426
447
405
394
530
1.5K
331

Last 14 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

2.04%18 Aug
2.60%
2.89%
2.69%
2.57%
3.06%
3.48%
2.82%
3.58%
3.46%
3.55%
3.40%
1.28%
3.02%

Reactions — likes, reposts, replies and quotes — divided by views. Median for under 10K accounts is 1.50%.

What the audience does

Likes61.5%790 in total
Reposts6.5%84 in total
Replies2.9%37 in total
Quotes2.3%29 in total
Bookmarks26.8%344 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Latest posts

@tomaarsen The blog post is amazing @tomaarsen you did a crazy work. I think this is a really good getting started with multi-vector models. It gives a strong overview 331 views · 9 likes · 0 reposts · 1 replies 18 Aug 2026 Antoine Chaffin, Raphaël Sourty & I wrote a blog post walking through all of this in practice: https://t.co/JulkQGDAJe pip install sentence-transformers==6.0.0 Full release notes: https://t.co/JQwqmYn96q 1.5K views · 17 likes · 1 reposts · 0 replies 18 Aug 2026 This release leans on a lot of other people's work. Thanks to @lateinteraction & @matei_zaharia for ColBERT, to @antoine_chaffin, @raphaelsrty, @paulomouraj & @AmelieTabatta at LightOn for PyLate & fast-plaid, and to Kisu Yang for the precision fixes. Full list in t 530 views · 17 likes · 0 reposts · 1 replies 18 Aug 2026 When I published v5.4, I wrote that it set up the groundwork for introducing late interaction models in the next major release. This is that release, and it's one of the largest updates in the project's history. 394 views · 13 likes · 0 reposts · 1 replies 18 Aug 2026 Also faster: multi-column losses now run one forward pass over merged columns, ~1.25x on hard-negative & triplet training, with identical loss trajectories. And fp16 + FlashAttention is now the fastest GPU config for SentenceTransformer at 3.87x over fp32. 405 views · 12 likes · 0 reposts · 2 replies 18 Aug 2026 Not multi-vector, but the fix I'd most want you to know about 🚨 CrossEncoder.predict now upcasts logits to float32 before the activation. A sigmoid in bfloat16 saturates & ties the top candidates together, randomizing their order. 0.1849 -> 0.6795 NanoBEIR nDCG@10. https 447 views · 15 likes · 0 reposts · 1 replies 18 Aug 2026 More 🚨 for v6.0: - transformers v5 & torch 2.2 are now the floor - similarity/similarity_pairwise are methods, not properties - custom module classes need trust_remote_code=True, local paths included - a bare list of chat messages is now one conversation, not a batch 426 views · 11 likes · 0 reposts · 1 replies 18 Aug 2026 Training works like the other model types: 4 new losses (incl. cached & distillation variants), 5 new evaluators, and a Trainer that takes the same arguments you already know. From a bare ModernBERT-base: 0.1338 -> 0.4831 NanoBEIR mean nDCG@10 in ~25 min on one RTX 3090. 489 views · 16 likes · 0 reposts · 1 replies 18 Aug 2026 51 checkpoints tested directly: 29 text retrieval & 22 visual document retrieval, from 17M to 8.8B parameters. LateOn, GTE-ModernColBERT, mxbai-edge-colbert, LFM2-ColBERT, colbertv2.0, answerai-colbert-small, ColPali, ColQwen, and more. 458 views · 13 likes · 0 reposts · 1 replies 18 Aug 2026 Sentence Transformers doesn't ship a late-interaction index, and doesn't need to: these indexes store whatever encode_document returned. Qdrant, Weaviate, Vespa, LanceDB, VectorChord & Milvus index multi-vectors natively, and LightOn's fast-plaid is a pip install away. 544 views · 13 likes · 0 reposts · 1 replies 18 Aug 2026 Because MaxSim is a sum of per-query-token maxima, a ranking decomposes exactly: every point of a score belongs to one query token & one document token. The new interpretability module renders that as the standard ColPali heatmap, aggregated or one map per query token. https 557 views · 14 likes · 0 reposts · 1 replies 18 Aug 2026 Page images aren't the only non-text modality. ColQwen-Omni takes text, images, audio & video. Retrieving a recorded conversation is the same two calls. Zero-shot, and no transcription step anywhere: the query says "nausea" where the audio says "carsickness". https://t.co/zc 588 views · 16 likes · 0 reposts · 1 replies 18 Aug 2026 Late interaction is the state of the art for visual document retrieval: text queries against page images, charts & tables intact, no OCR step. ColPali-family checkpoints run through the exact same two calls. MaxSim scores query text tokens against image patches. https://t.co 731 views · 16 likes · 1 reposts · 2 replies 18 Aug 2026 HierarchicalTokenPooling (Clavié, Chaffin & Adams) clusters each document's token vectors with Ward linkage & keeps ~1/pool_factor of them. pool_factor=2 halves the index at 100.6% of unpooled BEIR performance. Apply it per call, standalone, or bake it into the model. ht 686 views · 13 likes · 0 reposts · 1 replies 18 Aug 2026 The honest tradeoff is index size. 4,874 Natural Questions passages become 608,414 token vectors: 311.5 MB against ~20 MB for a simple 1024d dense index. About 16x. Three ways out: token pooling, a real late-interaction index, or using it as a reranker. 731 views · 15 likes · 0 reposts · 1 replies 18 Aug 2026 Does it actually help? LightOn trained LateOn (multi-vector) & DenseOn (dense) on the same data, same 149M ModernBERT backbone, differing only in whether they pool. Multi-vector wins 9 of 13 NanoBEIR datasets: 0.6868 vs 0.6764 mean nDCG@10. Same gap on full BEIR. https://t.c 1.3K views · 28 likes · 4 reposts · 1 replies 18 Aug 2026 Every checkpoint format loads through the same class: PyLate, Stanford-NLP ColBERT (via the HF_ColBERT marker + artifact.metadata), ColPali-style VLMs, or a bare backbone with a fresh projection. Prefixes, query expansion & the punctuation skiplist come from the saved config 2.5K views · 21 likes · 1 reposts · 1 replies 18 Aug 2026 Read MaxSim as a soft alignment: every query token points at the document token that best explains it. The alignment isn't lexical. In "Where do penguins live?" vs "Penguins inhabit Antarctica", `live` finds `inhabit` at 0.94, a word it shares no characters with. 1.1K views · 17 likes · 0 reposts · 1 replies 18 Aug 2026 Credit where it's due: ST handled dense & sparse but not late interaction, so LightOn built PyLate on top of it to close that gap. Much of what you can load today was trained with PyLate. With v6.0 those capabilities land in ST itself, designed together with its authors' wor 1.4K views · 22 likes · 2 reposts · 1 replies 18 Aug 2026 A dense model compresses a whole text into one vector. A multi-vector model keeps one vector per token & scores query against document with MaxSim: for each query token, take its best match in the document, then sum. Nothing has to be averaged away. https://t.co/lXi4HGjAmL 2K views · 29 likes · 0 reposts · 1 replies 18 Aug 2026 . @antoine_chaffin, @raphaelsrty & I wrote a blog post walking through multi-vector models in practice: loading checkpoints, scoring, search stacks, page images & keeping the index affordable. Or just point your Agent at the URL: https://t.co/JulkQGDAJe 2.8K views · 64 likes · 5 reposts · 1 replies 18 Aug 2026 🚨I've just released Sentence Transformers v6.0! MultiVectorEncoder joins the family: ColBERT-style late interaction models are now a first-class model type, for training, inference & interpretation, alongside dense, sparse & reranker models. Big thread 🧵 https://t.co/I 121.9K views · 399 likes · 70 reposts · 15 replies 18 Aug 2026

Similar accounts