Everybody assumed tokenization was already solved because the baseline was written in Rust and multithreaded. Marcel Røed spent a few days proving it was leaving three orders of magnitude on the table. HuggingFace tokenizers does 24.8 MB/s on GPT-2. GigaToken does 24.53 GB/s on an AMD EPYC 9565. Phi-4: 24.00 GB/s, 801×. Llama 3: 22.15 GB/s, 457×. On an M4 Max it’s 1,268×.
Where the speed comes from
It’s an MIT-licensed Rust library, not a service. The trick is pretokenization — the step normally handed to a regex engine. Røed rewrote it in SIMD, cached pretoken mappings aggressively, and cut branching and Python round-trips. HN gave it 274 points and GitHub gave it 1,000+ stars and 43 forks in the first few days. Rust keeps quietly eating AI infrastructure.
The API
pip install gigatoken. Two entry points: a compatibility mode that wraps your existing HuggingFace or tiktoken tokenizer with byte-identical output and a small speed tax, or the native GigaToken API where Rust reads files directly and parallelizes without Python in the loop. GPT-2, Llama, Qwen, DeepSeek, Gemma and Phi are covered. SentencePiece works but stays slow; WordPiece isn’t there yet.
If you run pretraining data pipelines, tokenization stops being a stage you budget for.
You Might Also Like
- Microgpt Andrej Karpathy Crammed an Entire gpt Into 243 Lines of Python and it Actually Works
- Phi 4 Reasoning Vision 15b Microsofts 15b Model Just Embarrassed gpt 4o on Vision Tasks
- Openai Swaps gpt 5 5 Instant in as Chatgpts Default Hallucinations Drop 52 on Legal and Medical Prompts
- The Open Interpreter Relaunch Rust Rewrite Codex Fork Bets Everything on Cheap Open Models
- Moltis a 60mb Rust Binary That Wants to be Your Entire ai Stack

Leave a comment