HeadlinesBriefing favicon HeadlinesBriefing.com

Meta's real-time audio model transcribes 20+ speakers

Engadget •
×

Meta has introduced its first real-time audio model, Muse Voice Transcribe, from Meta Superintelligence Lab (MSI). The model can handle dictation and transcription for more than 20 speakers and seamlessly handle multiple languages at once. Meta CEO Mark Zuckerberg shared an example demonstrating the model's ability to distinguish between speakers and switch between languages, including "code-switching."

"The model decides when to listen. It waits a little longer on hard words and commits faster on easy ones, using adaptive delay to predict each token and increase accuracy," Zuckerberg explained. "It holds up on messy, real audio too — trained across 70+ languages (with 25 validated at launch), handles mid-sentence code-switching, and manages hour-long sessions with 20+ speakers."

Meta's release follows Google Gemini 3.5 Transcribe by less than a week. While Google is integrating its model into Android and Chrome, Meta's plans for Muse Voice Transcribe remain unclear. Currently, the model powers dictation in Meta's Mac app and is available to developers via Muse Code and Meta's Model API, priced at $3 per 1,000 audio minutes.

This is MSL's latest release, following a coding agent and open-weight model. A demo is available on Meta's research blog.