Cohere opens a leading speech recognition model
Cohere released a 2-billion-parameter model under Apache 2.0 with an average word error rate of 5.42 per cent, taking first place on the open ASR leaderboard.
Why it matters
The most accurate dedicated speech recogniser became a 2-billion-parameter download under a permissive licence, moving transcription off the network and onto local hardware.
The model cohere-transcribe-03-2026 was published on 26 March 2026: 2 billion parameters, Apache 2.0 licence, trained on 14 languages - English, French, German, Italian, Spanish, Portuguese, Greek, Dutch, Polish, Mandarin Chinese, Japanese, Korean, Vietnamese and Arabic. Its average word error rate is 5.42 per cent, first place on the Hugging Face Open ASR Leaderboard. The published comparison gives Zoom Scribe v1 at 5.47, IBM Granite 4.0 1B Speech at 5.52, NVIDIA Canary Qwen 2.5B at 5.63 and OpenAI Whisper Large v3 at 7.44 per cent. The model is available as a download, through a free rate-limited API, or as a dedicated deployment. What the leaderboard says today. The Open ASR Leaderboard results file, rebuilt on 27 May 2026 on cleaned datasets, gives cohere-transcribe-03-2026 an average error of 4.67 per cent and places zoom/scribe_v1 at 4.15 and nvidia/canary-qwen-2.5b at 4.43 ahead of it; whisper-large-v3 stands at 5.78 there. The 5.42 and the first place belong to the leaderboard’s state on 26 March, which is no longer retrievable: until May its data lived in a private repository, and archived copies of the page are empty.