Back to timeline

Research · 1991

The vanishing gradient problem

Hochreiter showed formally why deep and recurrent networks do not train: the error signal decays exponentially with each layer backwards.

Why it matters

The main reason networks failed was explained, fifteen years before it was worked around.

On the backward pass the gradient is multiplied by activation derivatives smaller than one, so within a dozen steps it vanishes; with other values it explodes instead. The problem is therefore neither missing compute nor missing data but the arithmetic of the method itself. Written in German as a diploma thesis, the work went largely unnoticed by the English-language field; its consequences were acknowledged after LSTM and Bengio's 1994 paper.

Event record

Event date
1991
Timeline date
Event date
Verification
Sources gathered automatically · September 17, 2026
Lines
ID
evt-0164

The diploma thesis was completed in 1991 at the Technical University of Munich; the text is in German.

Sources

Related events

Records that link to this one