Research · September 2016
WaveNet
DeepMind trained a network to generate audio sample by sample, and synthesised speech first became indistinguishable from human speech by ear.
Why it matters
Speech synthesis moved to a generative model after forty years of filters and stitched-together fragments.
Earlier systems either assembled speech from recorded pieces or filtered an excitation by parameters, and both sounded mechanical. WaveNet models the distribution of each next sample given all previous ones, using dilated causal convolutions. The first version was too slow for practical use; an optimised version was running in Google Assistant a year later.
Event record
- Event date
- September 2016
- Timeline date
- Event date
- Verification
- Sources gathered automatically · September 17, 2026
- Lines
- ID
- evt-0232
Announcement of 8 September 2016.
Records that link to this one
- Related GPT-4o
Generating sound directly, without an intermediate text stage.
- Related Moshi
Generating audio directly with a neural network, eight years on.
- Contrasts Tacotron: speech straight from letters
The authors set themselves against WaveNet: slow because it generates sample by sample, and not end to end because it needs linguistic features from a separate frontend; Tacotron generates by frame and learns from text-audio pairs.
Tacotron: A Fully End-to-End Text-to-Speech Synthesis Model (arXiv:1703.10135v1) - Builds on Jukebox: songs with singing as raw audio
Jukebox's VQ-VAE encoders and decoders are built from "WaveNet-style" residual networks; the paper names WaveNet as the forerunner of autoregressive raw-audio generation.
Jukebox: A Generative Model for Music (arXiv:2005.00341v1) - Enables Unit selection: a voice from pieces of recording
WaveNet names this paper among the foundations of concatenative synthesis and takes a unit selection system as its baseline.
WaveNet: A Generative Model for Raw Audio (A. van den Oord, S. Dieleman, H. Zen et al.), arXiv:1609.03499v1, 12 September 2016 - Enables A voice from a statistical model: HTS
WaveNet describes statistical parametric synthesis with HMMs, citing Yoshimura's 2002 thesis, and takes a parametric system as a baseline.
WaveNet: A Generative Model for Raw Audio (A. van den Oord, S. Dieleman, H. Zen et al.), arXiv:1609.03499v1, 12 September 2016