Back to timeline

Research · October 13 – 16, 2003

Visual words: searching video like text

In October 2003 at ICCV Josef Sivic and Andrew Zisserman of Oxford carried the machinery of text retrieval over to images. Descriptors of regions in a frame were clustered into "visual words", the commonest and rarest were dropped by a stop list, frames were weighted by tf-idf and entered in an inverted file. Finding an outlined object among a film’s 4,000 keyframes took about 0.1 seconds.

Why it matters

Matching images stopped being a search over pairs of descriptors: matches are computed in advance, like words in documents, and an object is found across a whole film at once. The "bag of visual words" became the usual description of an image in recognition until convolutional networks: R-CNN in 2013 would still compare itself with such an approach.

Proceedings of the Ninth IEEE International Conference on Computer Vision, vol. 2, pp. 1470-1477. Two kinds of affine-invariant regions, about 1000 a frame on average, are tracked across neighbouring frames, and the 10% least stable tracks are rejected. The vocabulary was learnt by k-means on 48 shots of Run Lola Run, about 10,000 frames and 200,000 averaged descriptors: about 6,000 words for one kind of region and 10,000 for the other. The stop list removes the 5% commonest and 10% rarest words. The same vocabulary, unchanged, searched Groundhog Day. A query ran in Matlab on a 2 GHz Pentium.

Event record

Event date
October 13, 2003 – October 16, 2003
Timeline date
Event date
Verification
Sources gathered automatically · September 25, 2026
ID
evt-0816

The days of ICCV 2003 in Nice, per Crossref.

Sources

Related events

Records that link to this one