Jukebox: songs with singing as raw audio
On 30 April 2020 Prafulla Dhariwal and colleagues at OpenAI described Jukebox, a model that generates music with singing directly as a waveform: a multi-scale VQ-VAE compresses audio into discrete codes and Transformers model them. Generation can be steered by artist, genre and lyrics.
Why it matters
Music generation moved beyond notes and MIDI: the model produces coherent songs several minutes long with a voice, and the weights and code were published. The model is trained on recordings rather than scores.
The training set is 1.2 million songs, 600,000 of them in English, with lyrics and metadata from LyricWiki. Generation is slow: about an hour per minute of top-level codes and about eight more hours to upsample that minute to audio. The record does not claim the common "nine hours a minute" as the paper's figure: the paper gives the two parts separately.