Back to timeline

Research · 1996

Sounds found by how they sound

In the autumn of 1996 Erling Wold, Thom Blum, Douglas Keislar and James Wheaton of Muscle Fish in Berkeley described in IEEE MultiMedia a system that reduces any sound to loudness, pitch, brightness, bandwidth and harmonicity. With these features it retrieves similar sounds by example and assigns new ones to classes trained on examples, on a database of about 400 files.

Why it matters

Sound that is not speech, such as laughter, bells, a crowd or an animal, became searchable by content rather than by file name for the first time. The system explained its choices through features a person understands and borrowed the feature vector directly from image retrieval; classes here are defined by the user with a few examples.

The test database held about 400 sounds from sound effect and instrument sample libraries: animals, machines, instruments, speech and nature, from under a second to about a second and a half. A class is the mean feature vector of its examples and their covariance matrix; a new sound goes to the class at the smallest distance weighted by that covariance. On examples of laughter, female speech and telephone touch-tones the paper shows how the database is ordered by similarity. The analysis was already at work in Opcode Systems' Studio Vision Pro 3.0, which turned recordings into MIDI. What this record does not assert: the paper gives no share of correct answers, only ranked lists. Muscle Fish's 1995 papers (ICMC, an IJCAI workshop) are cited but were not opened. The paper does not cite QBIC.

Event record

Event date
1996
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0823

The Fall 1996 issue of IEEE MultiMedia (vol. 3, no. 3); no source prints a month. The IEEE Computer Society Digital Library's machine metadata gives 1 July 1996, the start of the quarter and earlier than 'Fall', so it is not taken as a date.

Sources

Related events

Records that link to this one