Qwen3-ASR: open weights for speech
On 29 January 2026 the Qwen team (Alibaba Cloud) released Qwen3-ASR in two sizes, 0.6B and 1.7B, for speech recognition and language identification across 52 languages and dialects, plus Qwen3-ForcedAligner-0.6B for aligning text with speech; per its report, under Apache 2.0.
Why it matters
Speech recognition across 52 languages and dialects became available under a licence that allows use and fine-tuning without an agreement with the developer. An editorial assessment; the model's quality is not assessed here, because the record has no independent measurement.
What the sources say. The developer's repository: released on 29 January 2026, two recognition models (0.6B and 1.7B) and the aligner Qwen3-ForcedAligner-0.6B; repository licence Apache-2.0; 52 languages and dialects for language identification and recognition, 11 languages for the aligner. Technical report arXiv:2601.21337 (v1 29 January, v2 30 January 2026): the models build on Qwen3-Omni, the aligner is non-autoregressive, the models are released under Apache 2.0. What the record does not claim. The developers' claims of quality (best result among open models, competitiveness with the strongest proprietary APIs, first-token latency, throughput under concurrency) sit in their own report and are not entered as facts: no independent measurement was found. The report calls only the aligner non-autoregressive, not the recognition itself as the outside report that named it described. The Hugging Face model pages, the blog post and the report's PDF were not opened.