Back to timeline

Research · September 5 – 9, 1999

A voice from a statistical model: HTS

At Eurospeech '99 in Budapest (5-9 September 1999) Takayoshi Yoshimura, Keiichi Tokuda and colleagues from the Nagoya and Tokyo Institutes of Technology described a synthesiser in which a hidden Markov model itself generates the spectrum, the pitch and the durations of sounds. On 25 December 2002 their group released the system openly as HTS 1.0.

Why it matters

A voice became the parameters of a model rather than a library of recordings: it can be trained on a few hundred sentences, stored compactly and changed without recording a new speaker. With unit selection it is one of the two lines of synthesis that WaveNet in 2016 names as its predecessors and baselines.

The system was trained on 450 phonetically balanced sentences of the ATR Japanese database (16 kHz, 25 mel-cepstral coefficients with deltas, three-state HMMs). Pitch is modelled with multi-space distributions because unvoiced sounds have none; spectrum, pitch and duration are clustered by three separate decision trees with 6,615, 1,877 and 201 leaves. Quality was checked only by informal listening; there are no figures. HTS 1.0 came out as a patch to HTK 3.2 under a free licence without commercial restrictions (once patched, HTK's licence applies), with no text analyser of its own; Festival played that part. The project's history page traces the line to a paper on parameter generation at ICASSP '95, which was not opened. What this record does not assert: the BSD licence appears in HTS's history only in 2008, for the hts_engine API, not for HTS 1.0.

Event record

Event date
September 5, 1999 – September 9, 1999
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
Lines
ID
evt-0824

The days of Eurospeech '99 in Budapest, per the paper's header; the open release HTS 1.0 is dated 25 December 2002 by its README.

Sources

Related events

Records that link to this one