The variational autoencoder
Kingma and Welling showed how to train a generative model with latent variables by gradient descent, by making randomness differentiable.
Why it matters
Probabilistic modelling and deep networks met in one method trainable by ordinary back-propagation.
The key device is reparameterisation: a random variable is written as a deterministic function of parameters and independent noise, and the gradient passes straight through. That removed the obstacle keeping probabilistic models from using deep networks. The method gives both generation and a compressed representation. Diffusion models belong to the same family of explicitly probabilistic approaches.