Unit selection: a voice from pieces of recording
At ICASSP-96 in Atlanta (7-10 May 1996) Andrew Hunt and Alan Black of the ATR laboratories in Kyoto described how the CHATR synthesiser builds speech from phonemes cut out of a large database of one speaker's recordings. Each phoneme is chosen by a Viterbi search on two costs: how close the unit is to the target and how smoothly it joins its neighbour.
Why it matters
Speech synthesis became a search through recorded data: the larger the database, the more natural the voice, and the cost weights are trained on natural speech instead of being tuned by hand. The approach held the best synthesis quality until neural models; WaveNet in 2016 names this paper among the foundations of the concatenative synthesis it measures itself against.
The units are phonemes. CHATR's databases ran from 10 to 150 minutes of speech, in Japanese and English, male and female. The target cost has 20-30 sub-costs (pitch, duration, power, phonetic context), the concatenation cost three: cepstral distance, and the differences in log power and pitch at the join; units adjacent in the recording join at no cost, so the search itself favours longer stretches. With a beam of 10-20 units the search runs near real time on a database of about 100,000 units (a Sun SPARCStation 20). The weights were trained in two ways: a search of the weight space (over 150 hours of computation on 40,000 units, about an hour of speech) and linear regression (1 to 10 hours). Both gave speech consistently better than hand-tuned weights; regression was only slightly better than the search. What this record does not assert: the paper gives no listening-test figures. Tacotron (2017) is compared with Google's later HMM-driven unit selection system, not with CHATR.