HeadlinesBriefing favicon HeadlinesBriefing.com

Canto: Speech Model for Real-World Dictation

Hacker News •
×

Speech recognition models excel at transcribing clean audio, but real dictation rarely happens under those conditions. Millions use Wispr Flow to message, write emails, code, and work through ideas at desks, between meetings, during commutes, and in busy offices. They speak through laptop microphones, earbuds, and headsets, often with other voices, music, or traffic in the background. Today, the Wispr Advanced Interfaces Lab is introducing Canto, our latest speech model for real-time dictation.

On an evaluation of real-world dictations, Canto achieved the lowest word error rate among all the models we tested. We compared Canto with models from Google, Open AI, Assembly AI, and Deepgram. Canto is the first model in a broader research and development program at Wispr Advanced Interfaces Lab. In this post, we share how it performs, how we trained it to handle challenging real-world conditions, and the research already shaping what comes next.

“Today, the Wispr Advanced Interfaces Lab is introducing Canto, our latest speech model for real-time dictation.” — Ariya Rastrow, CSO, Wispr Flow

To test Canto in real-world conditions, we created an evaluation set composed of 10 hours of English-language Wispr Flow dictations from more than 2,300 unique speakers, randomly sampled across applications and use cases. We enforced a strict separation between speakers in the train and test sets to avoid overfitting. Every sample came from a user who opted in to Wispr’s data-sharing setting, which allows their data to be used anonymously to evaluate and improve our models. More information about this setting is available in our data-controls documentation.

Canto achieved the lowest Word Error Rate (WER) of the models in our comparison. WER measures word substitutions, omissions, and insertions relative to a human-transcribed reference (lower is better).