Tacotron: speech straight from letters
On 29 March 2017 Yuxuan Wang and colleagues at Google described Tacotron, a model that synthesises speech directly from text characters and is trained from scratch on text-audio pairs. Its mean opinion score for naturalness was 3.82 of 5, against 3.69 for a production parametric system.
Why it matters
Speech synthesis, built from a text-analysis frontend, an acoustic model and a vocoder, became one network that can be trained without phonetic expertise. It opened the way to neural voices learned from recordings alone.
Training used about 24.6 hours of internal North American English; the waveform is recovered with the Griffin-Lim algorithm. A concatenative system scored 4.09 in the same test, so Tacotron did not reach it. Generating by frame makes the model substantially faster than methods that generate sample by sample.