Rectified linear units
Nair and Hinton showed that a simple nonlinearity which zeroes negative values trains better than a sigmoid.
Why it matters
One of the cheapest changes in the history of deep learning produced one of the largest speedups.
A sigmoid saturates: at large values its derivative is almost zero and the gradient decays with each layer. A rectified unit has derivative one across the whole positive half, so the signal reaches depth. It is also cheaper to compute than an exponential. Together with graphics processors this removed two of the three obstacles that had kept deep networks from working.