Sign language sentences from video
Thad Starner's master's thesis at the MIT Media Lab (February 1995, supervised by Alex Pentland) described a hidden Markov model system that recognises sentences of American Sign Language in real time from one camera's video, with a forty-word lexicon. In November, at ISCV in Coral Gables, they reported 99.2% word accuracy.
Why it matters
The method that had made speech recognition practical was carried to a language that is seen rather than heard: the machine read not single gestures but whole sentences from video. A sequence of hand movements proved fit for the same model as a sequence of sounds.
In the thesis (submitted on 20 January 1995) the hands wear coloured gloves, yellow on the right and orange on the left, tracked by one camera at 5 frames a second; 494 random five-word sentences (pronoun, verb, noun, adjective, pronoun) are split into 395 for training and 99 for testing; recognition runs five times faster than real time on an HP 735. On the independent test 97.0% of words were right with a grammar and 95.0% without, 90.7% accuracy counting insertions. The ISCV abstract gives 99.2% without modelling the fingers but names no conditions, and its full text was not read. The extended version appeared in IEEE PAMI in December 1998 (Starner, Weaver, Pentland): bare hands, the same forty-word lexicon; a desk camera gave 92%, a camera in the user's cap 98%, or 97% with an unrestricted grammar. The record does not claim the widely repeated 91.3% without grammar: the thesis gives 91% (90.7%).