VALL-E: a stranger's voice from three seconds
On 5 January 2023 Chengyi Wang and colleagues at Microsoft described VALL-E, a language model over discrete codes of the EnCodec neural codec, trained on 60,000 hours of English speech. A three-second recording of an unseen speaker is enough to synthesise any text in that voice.
Why it matters
Speech synthesis took up the recipe of large language models, scale of data and in-context learning, and voice cloning no longer required training recordings of the target speaker. The authors themselves named the risks: spoofing voice identification and impersonation.
The data are LibriLight, about 7,000 speakers. The word error rate of the synthesised speech is 5.9% against 7.7% for YourTTS, the best system then; 3.8% in continuation mode. The model keeps the emotion and acoustic setting of the prompt. The record does not claim the model was released: the paper does not say so.