Back to timeline

Research · September 2016

WaveNet

DeepMind trained a network to generate audio sample by sample, and synthesised speech first became indistinguishable from human speech by ear.

Why it matters

Speech synthesis moved to a generative model after forty years of filters and stitched-together fragments.

Earlier systems either assembled speech from recorded pieces or filtered an excitation by parameters, and both sounded mechanical. WaveNet models the distribution of each next sample given all previous ones, using dilated causal convolutions. The first version was too slow for practical use; an optimised version was running in Google Assistant a year later.

Event record

Event date
September 2016
Timeline date
Event date
Verification
Sources gathered automatically · September 17, 2026
Lines
ID
evt-0232

Announcement of 8 September 2016.

Sources

Related events

Records that link to this one