Informedia: searching video by what is said
The Informedia Digital Video Library at Carnegie Mellon, begun in 1994 under the Digital Library Initiative (NSF, DARPA, NASA), transcribed the sound of video with the Sphinx recogniser and found stories by the words of the transcript. By May 1998 it held over 1,000 hours of news and 400 hours of documentary video; the account appeared in IEEE Computer in February 1999.
Why it matters
Speech, image and text together described video so it could be searched by content rather than by the label on a tape. The paper measured the main point: search stayed useful even when the recogniser got nearly every third word wrong.
By the 1999 paper, first experiments on news in 1994 gave Sphinx a word error rate of 65%; Sphinx III trained on 66 hours of news gave about 24%, and a language model interpolated with the day's news from the CNN, Reuters and AP sites 19%. With error rates up to 30%, retrieval precision and recall fell by less than 10%. The library held over 40,000 segments, read captions off the frame and searched images by colour and regions. In September 1995 D-Lib Magazine described the News-on-Demand application: spoken queries, Sphinx-II on a separate machine, about 600 MB per hour of MPEG-I video. The record does not claim a date for the project's first publication, and gives its start by the paper with the measurements: another paper by the project's head says 1995.