Back to timeline

Milestone · October 1990

The TIMIT corpus

A joint effort by MIT, SRI and Texas Instruments produced a corpus of 630 speakers across eight dialects, annotated at the phoneme level.

Why it matters

Speech recognition got a shared ruler five years before text processing got the Penn Treebank.

Each speaker read ten phonetically rich sentences: 6300 sentences in all, more than five hours of speech with time-aligned transcription. The recording was done at Texas Instruments, transcription at MIT, verification and publication by NIST. The corpus is still used as a teaching example, though its size is long since trivial. What it established matters more: results became comparable between laboratories.

Event record

Event date
October 1990
Timeline date
Event date
Verification
Sources gathered automatically · September 17, 2026
Lines
ID
evt-0195

The full CD-ROM release was October 1990; a prototype disc appeared in December 1988 and the documentation in February 1993.

Sources

Related events

Records that link to this one