Meta Superintelligence Labs shipped its first real-time audio perception model, and it took the #1 spot on the Artificial Analysis streaming speech-to-text leaderboard on day one — while charging less than half what ElevenLabs does.
One model instead of three
Speech pipelines usually bolt together separate systems for transcription, speaker ID, and turn detection. Muse Voice Transcribe does all three in one model: streaming ASR, diarization for 20+ speakers, and endpointing. Audio comes in 80ms chunks, and the model decides how long to listen before committing each word. Trained on 70+ languages, it handles mid-sentence code-switching and stays coherent past an hour of conversation.
The numbers that matter
3.1% word error rate on AA-WER Streaming — ahead of Cartesia Ink-2 (3.4%), ElevenLabs Scribe v2 Realtime (3.6%), GPT Live Transcribe (3.9%), and Gemini 3.5 Transcribe Live (4.0%). Speaker diarization error is 17.5%, also first.
API access
Live on the Meta Model API at $3 per 1,000 audio minutes — about $0.18/hour, versus $6.50 for ElevenLabs. Streaming and batch endpoints cover meeting notes, voice agents, and call-center analytics. Also built into Meta AI for Mac and Muse Code. The catch: no open weights, API only.
You Might Also Like
- Openai gpt Realtime 2 Translate Whisper Three Voice Models one api Several Startups Erased
- Meta Model api Muse Spark 1 1 Undercuts Openai and Anthropic at 1 25 4 25 per Million Tokens
- Openai Realtime Voice Webrtc Stack the Infra Blueprint Every Voice Agent Startup now has to Compete With
- Google Gemini 3 5 pro ga Ships With a 2m Token Context Window the Biggest in any Production Model
- Openai gpt Realtime 2 1 gpt Realtime 2 1 Mini cut Voice Agent Latency by 25

Leave a comment