Sounds found by how they sound
In the autumn of 1996 Erling Wold, Thom Blum, Douglas Keislar and James Wheaton of Muscle Fish in Berkeley described in IEEE MultiMedia a system that reduces any sound to loudness, pitch, brightness, bandwidth and harmonicity. With these features it retrieves similar sounds by example and assigns new ones to classes trained on examples, on a database of about 400 files.
Why it matters
Sound that is not speech, such as laughter, bells, a crowd or an animal, became searchable by content rather than by file name for the first time. The system explained its choices through features a person understands and borrowed the feature vector directly from image retrieval; classes here are defined by the user with a few examples.
The test database held about 400 sounds from sound effect and instrument sample libraries: animals, machines, instruments, speech and nature, from under a second to about a second and a half. A class is the mean feature vector of its examples and their covariance matrix; a new sound goes to the class at the smallest distance weighted by that covariance. On examples of laughter, female speech and telephone touch-tones the paper shows how the database is ordered by similarity. The analysis was already at work in Opcode Systems' Studio Vision Pro 3.0, which turned recordings into MIDI. What this record does not assert: the paper gives no share of correct answers, only ranked lists. Muscle Fish's 1995 papers (ICMC, an IJCAI workshop) are cited but were not opened. The paper does not cite QBIC.