Speech model · Established · Intermediate
Whisper
An open speech recognition model family trained on large multilingual audio, widely used for transcription and translation.
What Whisper is
Whisper transcribes speech across many languages and can translate into English, with robustness to accents and background noise that made it a default open choice.
How it works
Weights are published and can be run locally in several optimised implementations, or accessed through hosted APIs. Smaller variants run in real time on modest hardware.
Why it matters
It made high-quality transcription free and self-hostable, which changed the economics of captioning and meeting tooling.
Common uses
- →Meeting and podcast transcription
- →Subtitling
- →Voice interfaces
- →Call analytics
Strengths
- ✓Open weights
- ✓Strong multilingual coverage
- ✓Runs locally
Watch for
- ✓Can hallucinate text in silence or noise
- ✓Diarisation needs a separate model
Continue exploring
More in this collection
Browse all AI Models