Local speech-to-text is fragmented. whisper.cpp only runs Whisper; Parakeet, Canary, Voxtral, and Kyutai each live in their own repos with their own quirks. Transcribe.cpp collapses all of them into one C/C++ inference library: 16+ ASR model families, 60+ models, all in GGUF format, all on the ggml runtime. Hacker News gave it 526 points.
The llama.cpp playbook, applied to voice
CJ Pais — author of Handy, the offline dictation app — built this as the first independent project out of Mozilla.ai’s Builders in Residence program, after shipping whisperfile and LocalScore for llamafile. The rigor shows: every model is numerically verified and WER-tested against its reference implementation, not just ported and hoped.
A library you can embed
It’s an MIT-licensed C/C++ library plus CLI, accelerated by Metal, Vulkan, CUDA, and TinyBLAS. Embed it and your app gets fully offline transcription — dictation, meeting notes, voice agents — with the freedom to swap Whisper for a faster Parakeet (the 110M model beats whisper base.en on speed) without touching your integration. llamafile already ships it as transcribefile.
One runtime, every model, runs everywhere: that bet built the llama.cpp ecosystem. Speech recognition just got its version.
You Might Also Like
- Ggml Llama cpp Joins Hugging Face and Honestly it was Only a Matter of Time
- Hypura Runs a 31gb Model on a 32gb mac at 2 2 tok s Llama cpp Just Ooms
- Cohere Transcribe Tops the Open asr Leaderboard With a 5 42 Word Error Rate and its Fully Open Source
- Ollama mlx on Apple Silicon 1810 Tokens sec Prefill and the end of Llama cpp on mac
- Mojo 1 0 Beta Lands After 3 Years one Codebase From cpu to gpu no Cuda

Leave a comment