Top AI Product

Every day, hundreds of new AI tools launch across Product Hunt, Hacker News, and GitHub. We dig through the noise so you don't have to — surfacing only the ones worth your attention with honest, no-fluff reviews. Explore our latest picks, deep dives, and curated collections to find your next favorite AI tool.


GigaToken hits 24.53 GB/s on GPT-2 — 989× faster than HuggingFace tokenizers

Everybody assumed tokenization was already solved because the baseline was written in Rust and multithreaded. Marcel Røed spent a few days proving it was leaving three orders of magnitude on the table. HuggingFace tokenizers does 24.8 MB/s on GPT-2. GigaToken does 24.53 GB/s on an AMD EPYC 9565. Phi-4: 24.00 GB/s, 801×. Llama 3: 22.15 GB/s, 457×. On an M4 Max it’s 1,268×.

Where the speed comes from

It’s an MIT-licensed Rust library, not a service. The trick is pretokenization — the step normally handed to a regex engine. Røed rewrote it in SIMD, cached pretoken mappings aggressively, and cut branching and Python round-trips. HN gave it 274 points and GitHub gave it 1,000+ stars and 43 forks in the first few days. Rust keeps quietly eating AI infrastructure.

The API

pip install gigatoken. Two entry points: a compatibility mode that wraps your existing HuggingFace or tiktoken tokenizer with byte-identical output and a small speed tax, or the native GigaToken API where Rust reads files directly and parallelizes without Python in the loop. GPT-2, Llama, Qwen, DeepSeek, Gemma and Phi are covered. SentencePiece works but stays slow; WordPiece isn’t there yet.

If you run pretraining data pipelines, tokenization stops being a stage you budget for.


You Might Also Like


Discover more from Top AI Product

Subscribe to get the latest posts sent to your email.



Leave a comment