The vanishing gradient problem
Hochreiter showed formally why deep and recurrent networks do not train: the error signal decays exponentially with each layer backwards.
Why it matters
The main reason networks failed was explained, fifteen years before it was worked around.
On the backward pass the gradient is multiplied by activation derivatives smaller than one, so within a dozen steps it vanishes; with other values it explodes instead. The problem is therefore neither missing compute nor missing data but the arithmetic of the method itself. Written in German as a diploma thesis, the work went largely unnoticed by the English-language field; its consequences were acknowledged after LSTM and Bengio's 1994 paper.