FaceNet: a face as 128 bytes
On 12 March 2015 Google described FaceNet: a network turns a face into a 128-byte embedding in which distance means similarity. On Labeled Faces in the Wild, 99.63%; on YouTube Faces DB, 95.12%; on both the error is 30% below the best published result.
Why it matters
Recognition, verification and clustering of faces became distances in a compact space trained directly with same-or-different triplets. Such embeddings can be compared at scale without retraining for new faces.
Trained on 100-200 million face thumbnails of about 8 million identities; triplets are mined during training. The record does not claim how or where Google deployed the system: the paper does not say.