pix2pix: one model for image translation
On 21 November 2016 Phillip Isola, Alexei Efros and colleagues at Berkeley showed that a conditional adversarial network with one architecture and objective translates image to image across many tasks: labels to street scenes, aerial photos to maps, day to night, edges to photos.
Why it matters
Tasks that had each needed their own algorithm became one model trained on different pairs. The loss for each task no longer had to be designed by hand either: the discriminator learns it.
The generator is a U-Net type, and a PatchGAN discriminator judges separate patches of the image. In a study on Amazon Mechanical Turk, aerial photos generated from maps fooled participants on 18.9% of trials, well above the L1 baseline.