Back to timeline

Research · March 29, 2017

Tacotron: speech straight from letters

On 29 March 2017 Yuxuan Wang and colleagues at Google described Tacotron, a model that synthesises speech directly from text characters and is trained from scratch on text-audio pairs. Its mean opinion score for naturalness was 3.82 of 5, against 3.69 for a production parametric system.

Why it matters

Speech synthesis, built from a text-analysis frontend, an acoustic model and a vocoder, became one network that can be trained without phonetic expertise. It opened the way to neural voices learned from recordings alone.

Training used about 24.6 hours of internal North American English; the waveform is recovered with the Griffin-Lim algorithm. A concatenative system scored 4.09 in the same test, so Tacotron did not reach it. Generating by frame makes the model substantially faster than methods that generate sample by sample.

Event record

Event date
March 29, 2017
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0772

The day the first version of the preprint was submitted; the figures were read in it. The second version of 6 April 2017 renamed it "Tacotron: Towards End-to-End Speech Synthesis".

Sources

Related events

Records that link to this one