Back to timeline

Research · April 30, 2020

Jukebox: songs with singing as raw audio

On 30 April 2020 Prafulla Dhariwal and colleagues at OpenAI described Jukebox, a model that generates music with singing directly as a waveform: a multi-scale VQ-VAE compresses audio into discrete codes and Transformers model them. Generation can be steered by artist, genre and lyrics.

Why it matters

Music generation moved beyond notes and MIDI: the model produces coherent songs several minutes long with a voice, and the weights and code were published. The model is trained on recordings rather than scores.

The training set is 1.2 million songs, 600,000 of them in English, with lyrics and metadata from LyricWiki. Generation is slow: about an hour per minute of top-level codes and about eight more hours to upsample that minute to audio. The record does not claim the common "nine hours a minute" as the paper's figure: the paper gives the two parts separately.

Event record

Event date
April 30, 2020
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0773

The day the first version of the preprint was submitted; the figures were read in it.

Sources

Related events

Records that link to this one