HeadlinesBriefing HeadlinesBriefing.com

Whistle: Open Speech to Text in 16.9 MB

Hacker News •
×

Today we release Whistle, an open speech recognition model built for mobiles, wearables, robots, smart home devices, automotive systems and microcontrollers. It is a single 16.9 MB file that runs on the CPU with no dependencies. It loads into the same C++ engine as Needle, sharing its container and quantisation, so one binary can turn a spoken clip directly into tool calls.

Whistle does three jobs entirely on the device. It transcribes up to 30 seconds of 16 kHz mono audio in English, German, French, Spanish, Italian, Dutch and Polish, detecting the language unless one is named. It produces word timestamps with a start, end and probability for every word, aligned from the decoder's attention. It can also output a speech embedding, one row per 80 ms frame, without decoding a transcript. The model reaches the first token in 11 ms.

The architecture has two halves. The front end converts audio into 80 log-mel bins and uses a convolutional stem to reduce 30 seconds to 375 frames. The encoder then applies eight Simple Attention blocks shared with Needle, with attention that is not causal, so a frame at 3 seconds can attend to one at 12 seconds. The decoder uses eight Laddered Simple Attention blocks at width 512, with gated cross attention in every layer that reads the encoder output. Those projections are computed once per clip, so five beam search paths cost only five short transcript caches rather than five passes over the audio.

Decoding uses five beams scored by length-normalised log probability. Keyword biasing runs an Aho-Corasick automaton over phrases supplied by the user, raising their probability as the search advances. The vocabulary holds 8,192 text pieces plus seven language tokens, and transcripts are capped at 320 tokens.

The model is also available in a browser sandbox where audio never leaves the device.

Source: Hacker News · Summarized by HeadlinesBriefing