Back to timeline

Benchmark · September 4 – 8, 2005

Blizzard Challenge: speech synthesis on shared data

At Interspeech 2005 in Lisbon (4-8 September 2005) Alan Black (Carnegie Mellon) and Keiichi Tokuda (Nagoya Institute of Technology) reported the first contest of speech synthesisers on shared data: six teams from three continents built voices in a week or two from the same CMU ARCTIC databases and spoke the same 250 sentences. Listeners rated naturalness and typed in what they heard; the HMM-based system from Nagoya came first with every group of listeners.

Why it matters

Speech synthesis gained what recognition had had since the 1980s: shared data, the same tests and comparable numbers instead of each laboratory demonstrating on its own recordings. The speakers' own voices, recorded alongside, showed the distance: no system came close to them.

The CMU ARCTIC databases hold about 1200 sentences per voice from Project Gutenberg books, at 16 kHz; two voices had been available since June 2004, two more came with the test sentences on 15 January 2005. Five genres: novels, news and conversational phrases were rated on a scale of 1 to 5, while rhyming words and semantically unpredictable sentences were typed in by listeners and scored as the share of words wrong. The listeners were experts from each team, volunteers from the web and paid undergraduates. With the experts, natural speech scored 4.76 and 8.5% word errors, the best system 3.19 and 14.7%, the weakest about 2.1 and 32.7%. The report hides the teams behind the letters A to F, but three of them described themselves. Nagoya's HTS system, built on hidden Markov models, had by its own paper the highest scores and the fewest errors with every group of listeners, and in the report's table only system D is first in every column, so D is Nagoya's. IBM's concatenative system, designed for much larger databases, was second on scores by its own paper; CMU's Festival unit selection was system C, "in the middle of the pack". The Nagoya group puts the lead down to the data: on about an hour of speech a voice, HMMs cover it better, while unit selection usually needs more than 10 hours. What this record does not assert: no document read says who stands behind the other letters.

Event record

Event date
September 4, 2005 – September 8, 2005
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0829

The days of Interspeech 2005 in Lisbon per the ISCA Archive; participants received the test sentences on 15 January 2005.

Sources

Related events