Whistle: Speech to Text in 16.9 MB

826 points · 167 comments on HN · read original →

Points and comments are a snapshot, not live.

Whistle is a 16.9 MB speech-to-text model that runs on-device with CPU and no dependencies.

Whistle transcribes 7 languages (English, German, French, Spanish, Italian, Dutch, Polish) in one pass up to 30 seconds. It reaches first token in 11 ms on an Apple M4 Pro CPU, decodes at 1,319 tokens/s, and uses a 5-beam search. The model uses 8 Simple Attention encoder blocks and an 8-layer Laddered Simple Attention decoder. Word error rates are lower than Whisper base on LibriSpeech and SPGISpeech, but Whisper leads on TED-LIUM, AMI, and MLS average. The same C++ engine can load both Whistle and Needle models, enabling speech-to-tool pipelines without passing transcripts externally.

The model is available via pip (`cactus-needle`) and as prebuilt binaries for 17 platforms including iOS, Android, and the browser. Weights are on Hugging Face.

What commenters are saying

English accuracy is strong, even with non-native accents, but Spanish transcription produces typos and non-existent words for Mexican and Venezuelan speech. Several commenters note the real challenge is not model size but understanding atypical speech patterns, such as elderly or stroke-impaired speakers. Recommendations include using dedicated dictation models that ignore filler words and pairing transcription with a cheap LLM cleanup pass. Handy is cited as a useful tool for streaming dictation, and Gemini Desktop is praised for handling heavy accents.