tweetindex
ES

Marcel Rød

@marcelroed

PhD Student at Stanford working with Tatsu Hashimoto and Jure Leskovec. Previously MIT, Oxford, CERN

2 811Followers
281Following
31Posts total
1.1MViews on collected posts

Últimas publicaciones

@marcelroed Very impressive!! Although you've apparently far exceeded it and I am sure you hadn't heard of it anyways, I feel compelled to plant my flag here (for history) that was already well past the two comparisons in your chart over a year ago, at ~1.2 GB/s (and 2 GB/s+ mor 19.7K views · 160 likes · 2 reposts · 2 replies Open on X →
I had a lot of fun finding these optimizations over the past year! Thanks to @mdbereket, @percyliang, @tatsu_hashimoto, @jure for feedback and supervision. Check out Gigatoken on GitHub! https://t.co/IPme9IeFDv 17.9K views · 267 likes · 8 reposts · 4 replies Open on X →
If you're curious about how fast Gigatoken will run on your data and with your tokenizer, you can run it without installing Gigatoken using uv: uvx gigatoken bench <tokenizer-name> <dataset> https://t.co/dEA1Ftg0iZ 18.7K views · 155 likes · 0 reposts · 1 replies Open on X →
Gigatoken can be used as a drop-in replacement for HuggingFace tokenizers or tiktoken in Python, or with its own API. https://t.co/IpMdiLWZk3 19.2K views · 188 likes · 1 reposts · 1 replies Open on X →
Caching is the other main optimization, where Gigatoken avoids recomputing mappings for pretokens it’s already seen. A fast cache hierarchy lets this be much faster than recomputing, where other implementations (like tiktoken) report worse performance with caching. 20.3K views · 172 likes · 1 reposts · 1 replies Open on X →
After writing my own tokenizer a few times for our class, https://t.co/Aw4AV3TKqQ, I tried moving the pretokenizer (pre-splitting documents into words) from a regex engine to a hand-written state machine, and realized this could also be done using CPU-specific SIMD instructions. 24.6K views · 252 likes · 6 reposts · 6 replies Open on X →
Introducing the world's fastest tokenizer implementation, Gigatoken! Gigatoken is ~500-1000x faster than HuggingFace, and ~100x faster than OpenAI's tiktoken for most tokenizer definitions on most machines. These baselines are already multithreaded Rust implementations! 🧵 https 734.4K views · 4.1K likes · 483 reposts · 85 replies Open on X →
Inspired by @lemire's simdjson, as well as @cmuratori's Handmade software movement, I noticed that existing tokenizers were running at much slower than the data rates I've seen for programs doing a comparable amount of work. For instance, simdjson processes several GB/s. 27.6K views · 217 likes · 1 reposts · 2 replies Open on X →
Gigatoken can tokenize the entire internet in <7 hours on one machine. Specifically, this refers to tokenizing Common Crawl on a dual socket AMD Epyc server, reaching 24GB/s of throughput, as seen here: https://t.co/jroCWjxKxi 43.9K views · 550 likes · 41 reposts · 6 replies Open on X →
We built the fastest BPE tokenizer in the world. 100% correct on models like llama3/gpt-4o/inversion—over 100x faster than other libraries—to power the world's fastest language models in Inversion 2. https://t.co/6Vaua4cGhI 143.8K views · 1.1K likes · 64 reposts · 30 replies Open on X →

Frente a cuentas del mismo tamaño

9 publicaciones de los últimos 90 días, junto al rango de under 10K seguidores. llega a mucha gente, pero pocos de esos espectadores reaccionan.

Visualizaciones medianas20 320esta cuenta3 378mediana de under 10K
Alcance, %722.87%esta cuenta247.46%mediana de under 10K
Interacción, %0.86%esta cuenta1.44%mediana de under 10K
MétricaEsta cuentaMediana de under 10KProporción
Visualizaciones medianas por publicación20 3203 3786.02×
Alcance (visualizaciones ÷ seguidores)7.2× audience2.5× audience2.92×
Tasa de interacción0.86%1.44%0.59×

Otras cuentas de este rango →   Comparar con otra cuenta →   Cómo se construyen estas referencias →

Growth & engagement

How the posts we collected actually performed: views and reaction rate post by post, what the audience did with them, and where the follower count goes.

Views per post

143.8K25 Feb
43.9K21 Jul
27.6K
734.4K
24.6K
20.3K
19.2K
18.7K
17.9K
19.7K

Last 10 collected posts, oldest on the left. The scale is logarithmic: one post can outrun the rest a hundred times over.

Engagement rate per post

0.84%25 Feb
1.37%21 Jul
0.80%
0.65%
1.07%
0.86%
0.99%
0.83%
1.56%
0.83%

Reactions — likes, reposts, replies and quotes — divided by views. Median for under 10K accounts is 1.44%.

What the audience does

Likes62.7%7 204 in total
Reposts5.3%607 in total
Replies1.2%138 in total
Quotes0.9%107 in total
Bookmarks29.9%3 442 in total

Share of every reaction we collected for this account. Replies mean argument, reposts mean endorsement, bookmarks mean the post was worth keeping.

The follower curve appears once this account has two daily snapshots — we take one a day, and this one is on its first.

Cuentas similares