"Put-That-There": voice and gesture in one command
In July 1980 Richard Bolt of MIT's Architecture Machine Group described at SIGGRAPH a system in which a person seated before a wall-sized screen creates and moves shapes by saying "put that there" and pointing. An NEC DP-100 with a vocabulary of up to 120 words recognised the speech, and a Polhemus magnetic sensor on a watchband gave the direction of the arm.
Why it matters
Two input channels work as one sentence: "that" and "there" get their meaning only from the gesture, and the gesture becomes precise through the word. The problem later taken up by Kinect and multimodal assistants is posed here and shown as a working system.
By the paper, the DP-100 took up to five words a sentence without pauses and answered about 300 ms after its end; in discrete-word mode the vocabulary reached about 1,000. The ROPAMS sensor from Polhemus Navigation Sciences is a 1.5-inch transmitter cube and a 0.75-inch sensor cube in a nutating magnetic field. The Media Room is about 16 by 11 feet. Commands: create, move (with the synonyms put and place), copy, make that, delete, name. DARPA funded the work. The paper describes a system, not an experiment: the record claims neither a recognition accuracy nor any measure of how people fared using it.