Visual words: searching video like text
In October 2003 at ICCV Josef Sivic and Andrew Zisserman of Oxford carried the machinery of text retrieval over to images. Descriptors of regions in a frame were clustered into "visual words", the commonest and rarest were dropped by a stop list, frames were weighted by tf-idf and entered in an inverted file. Finding an outlined object among a film’s 4,000 keyframes took about 0.1 seconds.
Why it matters
Matching images stopped being a search over pairs of descriptors: matches are computed in advance, like words in documents, and an object is found across a whole film at once. The "bag of visual words" became the usual description of an image in recognition until convolutional networks: R-CNN in 2013 would still compare itself with such an approach.
Proceedings of the Ninth IEEE International Conference on Computer Vision, vol. 2, pp. 1470-1477. Two kinds of affine-invariant regions, about 1000 a frame on average, are tracked across neighbouring frames, and the 10% least stable tracks are rejected. The vocabulary was learnt by k-means on 48 shots of Run Lola Run, about 10,000 frames and 200,000 averaged descriptors: about 6,000 words for one kind of region and 10,000 for the other. The stop list removes the 5% commonest and 10% rarest words. The same vocabulary, unchanged, searched Groundhog Day. A query ran in Matlab on a 2 GHz Pentium.