MusicLM: music from a text description
On 26 January 2023 Andrea Agostinelli and colleagues at Google described MusicLM, a model that generates music from a description such as "a calming violin melody backed by a distorted guitar riff": 24 kHz, coherent over several minutes. A melody can also be hummed or whistled, and the model plays it in the described style.
Why it matters
Text became a way to steer music as it already steered images. With the paper Google released MusicCaps, the first evaluation set collected for text-to-music generation, but not the model itself.
MusicLM rests on AudioLM and on MuLan's joint space of music and text; training used five million clips, 280,000 hours of music. MusicCaps is 5,521 music-caption pairs written by musicians. A memorisation check found exact matches for a tiny fraction of examples and approximate ones for 1%; the authors wrote that they had "no plans to release models at this point".