Faces from the news as a recognition benchmark
In October 2007 Gary Huang, Manu Ramesh, Tamara Berg and Erik Learned-Miller of the University of Massachusetts Amherst released Labeled Faces in the Wild: 13,233 photographs of the faces of 5749 people from news articles on the web, 1680 of them in two or more photos. The task is to say of a pair of photos whether they show the same person.
Why it matters
Face databases had until then been shot under controlled conditions, with pose, lighting and background set. LFW took faces as the press photographs them, with a single filter, the Viola-Jones detector, and a fixed protocol of ten splits. Face recognition was measured on it for the next decade: DeepFace in 2014 would approach human level on it, FaceNet in 2015 would pass 99%.
Technical Report 07-49. Images are 250 × 250 pixels, each naming the person at the centre of the frame; 4069 people have a single photo. The basis was the "Faces in the Wild" set of Berg and colleagues, built by analysing news photos together with their captions, with 77% label accuracy; LFW removed its errors and duplicates. The first view of the data has 1100 pairs of the same person and 1100 of different people for training and 500 of each for testing; the second, ten subsets for cross-validation. The authors warn of the source’s biases: few photos in poor light, and because of the Viola-Jones detector few profiles or views from above or below. The record reads an early revision of the report from Learned-Miller’s page; a later one, archived from the dataset site, differs in its references, not its figures. The dataset site itself no longer opens.