Shape and motion from a rank-three matrix
In November 1992 Carlo Tomasi and Takeo Kanade showed that if the coordinates of P points tracked through F frames are collected into one 2F by P matrix, then under orthographic projection and without noise that matrix has rank exactly three. A singular value decomposition splits it into camera motion and scene shape — one computation over the whole sequence, with no error accumulating frame to frame.
Why it matters
Recovering three-dimensional shape from video stopped being a chain of estimates between neighbouring frames. The whole trajectory goes into one matrix, and a property of that matrix rather than the quality of any single step decides the solution; there is a straight line from here to the batch optimisation in modern mapping systems.
The laboratory experiment: the "Hotel" stream of 150 frames, shot with a Sony camera and a 200 mm lens on a high-precision positioning platform. Automatic selection produced 430 features, 42 of which were abandoned during tracking, and the trajectories of the remaining 388 formed the measurement matrix. Rotation errors against the platform readings are everywhere less than 0.4 degrees and on average 0.2 degrees. The second experiment is about occlusion: a fill matrix of 226 by 829 holds 187,354 cells of which 30,185, about 16 percent, are known. The method reconstructs the remaining 84 percent. The record does not claim "600 frames", "100 feature points" or "error under 1 percent of scene depth". None of those figures is in the paper. At publication Tomasi is in the Department of Computer Science at Cornell University and Kanade in the School of Computer Science at Carnegie Mellon.