Dropout
On 3 July 2012 Geoffrey Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever and Ruslan Salakhutdinov of the University of Toronto posted a preprint: on every training case, omit each hidden unit at random with probability 0.5, and halve the outgoing weights at test time. On MNIST the best standard network made 160 test errors; dropout brought it to about 130, and dropping 20 per cent of the input pixels as well to about 110. A single ImageNet 2010 network went from 48.6 to 42.4 per cent error.
Why it matters
A large network on limited data could be kept from overfitting without training many separate models: dropout trains an exponential number of thinned networks that share weights and averages them cheaply at test time. The AlexNet paper that autumn used it in its first two fully connected layers and said that without it the network overfits substantially.
The preprint reports records in speech and object recognition. TIMIT frame classification: error from 22.7 to 19.7 per cent, a record for methods that use no information about speaker identity. CIFAR-10: best published result without transformed data 18.5 per cent, their convolutional network 16.6, with dropout in the last hidden layer 15.6. ImageNet 2010: winning entry 47.2 per cent, state of the art 45.7, their single network 48.6, with 50 per cent dropout in the sixth hidden layer 42.4. Instead of an L2 penalty the authors cap the norm of each unit's incoming weights. Two dates: the preprint of 3 July 2012 and the JMLR article (submitted November 2013, published June 2014).