An object as a template with movable parts
At CVPR in June 2008 Pedro Felzenszwalb (University of Chicago), David McAllester (TTI Chicago) and Deva Ramanan (UC Irvine) described a detector in which an object is a coarse template on gradient histograms plus finer part templates that can shift. Trained with a latent SVM, it doubled the best PASCAL 2006 result for people: 0.34 average precision against 0.16.
Why it matters
Part models had long been an appealing idea but lost to simpler rigid templates on hard datasets. This work showed their advantage on PASCAL, beating the best results of the 2007 challenge in ten categories of twenty, and became the model new work was compared against: R-CNN in 2013 would call it "venerable" and measure itself against it.
Proceedings of CVPR 2008. Part positions in the training examples are not labelled: they are latent, the latent SVM chooses them itself, and training becomes convex once they are fixed for the positive examples. Hard negative examples are mined from false detections. On the PASCAL 2006 person set: the challenge winner 0.16, the best previous result 0.19, the coarse template alone 0.18, with latent choice of position 0.24, with parts 0.34. Training takes 3-4 hours a class, processing an image about 2 seconds.