R-CNN: a convolutional network finds objects
On 11 November 2013 Ross Girshick, Jitendra Malik and colleagues at Berkeley showed that a network trained to classify whole ImageNet images could find objects: it scores about 2,000 region proposals, and mean average precision on PASCAL VOC 2007 reaches 48%.
Why it matters
Object detection moved from hand-made features to convolutional features transferred from classification, gaining over 40% relative to deformable part models. The line of detectors that measured themselves against R-CNN starts here.
The pipeline: bottom-up region proposals (selective search), features from a large convolutional network for each, then linear SVMs per class. On PASCAL VOC 2010, 43.5% mAP against 35.1% for a method using the same regions with a bag of visual words and 29.6% for deformable part models. The record does not claim 53.3% on VOC 2012 or an ILSVRC2013 result: those figures are not in the first version and were added in later ones.