Benchmark · 1988
Sphinx
Kai-Fu Lee's system was the first to recognise continuous, large-vocabulary speech from any speaker, with no tuning to the voice.
Why it matters
Three requirements considered incompatible were met together, and that combination is what made recognition usable.
Before Sphinx a system either required pauses between words, or worked with a small vocabulary, or had to be trained on a particular speaker. Lee combined hidden Markov models with context-dependent triphones and parameter sharing between them, which let the model train on limited data. Business Week named it the most important innovation of 1988. The open-source CMU Sphinx line exists to this day.
Event record
- Event date
- 1988
- Timeline date
- Event date
- Verification
- Sources gathered automatically · September 17, 2026
- Lines
- ID
- evt-0193
The thesis was completed in 1988 at Carnegie Mellon and described in the February 1989 workshop proceedings.
Records that link to this one
- Related The tutorial account of hidden Markov models
It describes the apparatus the systems of the same period are built on.
- Related The TIMIT corpus
It supplies the shared data that systems of that generation began to be measured on.
- Builds on Informedia: searching video by what is said
The video transcripts came from Carnegie Mellon's Sphinx recogniser: Sphinx-II in the 1995 paper, Sphinx III in the 1999 one.
Lessons Learned from Building a Terabyte Digital Video Library (course copy, UC Berkeley School of Information, i246)News-on-Demand: An Application of Informedia Technology (Alexander G. Hauptmann, Michael J. Witbrock, Michael G. Christel), D-Lib Magazine, September 1995 - Related Searching recorded speech: the SDR track
NIST first took SPHINX-III as its baseline recogniser, but it ran at nearly 200 times real time, and a faster one had to be found for 87 hours.
The TREC Spoken Document Retrieval Track: A Success Story (John S. Garofolo, Cedric G. P. Auzanne, Ellen M. Voorhees), RIAO-2000 Content-Based Multimedia Information Access, Paris, 12-14 April 2000 - Related The TREC video track: a shared test for video search
Carnegie Mellon searched the video both with and without the Sphinx speech recogniser: the track's rules required a run without transcripts.
The TREC-2001 Video Track Report (Alan F. Smeaton, Paul Over, R. Taban), in NIST Special Publication 500-250, The Tenth Text REtrieval Conference (TREC 2001); report dated 18 April 2002 - Enables Resource Management: a shared speech recognition test
Sphinx was measured on exactly this task: 96 percent of words with the word-pair grammar on the official test sentences of March and October 1987.
Recent Progress in the SPHINX Speech Recognition System