Musical genre recognised from sound
In July 2002 George Tzanetakis and Perry Cook of Princeton University published in IEEE Transactions on Speech and Audio Processing a system that assigns a music recording to a genre from features of timbre, rhythm and pitch. On a set of ten genres with a hundred 30-second excerpts each it was right 61% of the time, where chance would give 10%.
Why it matters
Genre, a human label without sharp boundaries, was measured for the first time as a machine recognition task with a shared set and an accuracy figure. These ten genres of a hundred excerpts became, as GTZAN, the most-used public set for genre recognition: a 2013 review found it in at least a hundred works, and with it repetitions, mislabellings and distortions that almost no one had noticed.
For each of 20 music genres and three speech genres 100 excerpts of 30 seconds were collected, about 19 hours from radio, CDs and MP3s, in 22,050 Hz mono. The ten-genre set: classical, country, disco, hip-hop, jazz, rock, blues, reggae, pop, metal. Each recording is described by 30 numbers: 19 for timbre, 6 for rhythm (a beat histogram), 5 for pitch; the classifiers are Gaussian mixtures and k nearest neighbours. The system told music from speech 86% of the time and three kinds of speech 74%. It was built on the free MARSYAS framework under the GNU GPL. What this record does not assert: the name GTZAN is not in the paper, and this record does not say where it came from; according to Tzanetakis, quoted in the same 2013 review, the set was not made expressly for genre recognition. The comparison with people is taken from another study with different genres: students were right 53% of the time after a quarter of a second and 70% after three seconds.