Apple’s SpeechAnalyzer just beat Whisper a few months ago. Now Google fires back. Gemini 3.5 Transcribe launched August 26 and immediately hit both Product Hunt and the HackerNews front page (181 points).
It’s a dedicated speech-to-text model, not a chatbot feature: 2.6% word error rate on recorded audio, 4.0% streaming, 70% lower latency than its predecessor Chirp 3. 85+ languages with mid-sentence code-switching, speaker diarization, word-level timestamps.
Transcription that thinks
The actual pitch isn’t accuracy — it’s reasoning. The model strips filler words, handles self-corrections (“meet at 3, no wait, 4” becomes “meet at 4”), and formats output automatically. Whisper gives you what you said; this gives you what you meant. It already powers Gboard’s Rambler feature and the Gemini app on macOS, with Chrome next.
The API
Two endpoints through the Gemini API: gemini-3.5-transcribe-live for real-time streaming (voice agents, live captions) and gemini-3.5-transcribe for recorded audio with speaker labels (meeting notes, call analytics). Custom vocabulary biasing handles domain jargon. Roughly $0.005/min batch, $0.009/min live — free in AI Studio preview. Self-hosted Whisper stays cheaper for English-only batch jobs. For everything else, Google just reset the bar.
You Might Also Like
- Apple Speechanalyzer api Benchmark 55 Faster Than Whisper and More Accurate in English
- Starnus Just hit 1 on Product Hunt and Yeah its Worth the Hype
- Lovon Just Topped Product Hunt on Valentines day and its not a Dating app
- Zenmux Just hit 1 on Product Hunt Heres why Everyones Paying Attention
- Google Lyria 3 Just Turned Gemini Into a Music Studio and im Weirdly Into it

Leave a comment