Gradient histograms find pedestrians
In June 2005 Navneet Dalal and Bill Triggs showed that a dense grid of histograms of oriented gradients, normalised over overlapping blocks and fed to a linear support vector machine, detects people in images at least ten times more accurately than existing feature sets. Alongside the method they released the INRIA Person set of 1805 human images at 64 by 128 pixels.
Why it matters
Detecting objects of variable shape acquired a feature good for something harder than a face. The Viola-Jones method worked on faces because faces are alike; a person in an arbitrary pose does not yield to that, and HOG held this position until convolutional detectors arrived.
The settings the paper calls its default: 9 orientation bins over 0 to 180 degrees, 16 by 16 pixel blocks of four 8 by 8 cells, a 64 by 128 detection window, a block stride of 8 pixels, L2-Hys normalisation. The best result in the sweep over geometries comes from different sizes: 3 by 3 cell blocks of 6 by 6 pixel cells give a 10.4 percent miss rate at 10 to the minus 4 false positives per window. The record does not attribute 10.4 percent to the default setting: these are two separate lines in the paper, and the figure belongs to 3 by 3 blocks of 6 by 6 cells, not to 16 by 16 blocks of 8 by 8 cells. The INRIA Person set holds 1805 images, not 1807. It was made because on the older MIT set of 509 training and 200 test images the detector gave essentially perfect separation — that is, the set had stopped measuring anything. The paper puts the comparison with other features this way: HOG-based detectors give at least an order of magnitude reduction in false positives per window on the INRIA set against the authors' own implementations of Haar wavelets, PCA-SIFT and shape contexts.