Why long-term dependencies are not learned
Bengio, Simard and Frasconi proved that gradient training of recurrent networks cannot both hold information for a long time and stay robust to noise.
Why it matters
The problem was stated as a trade-off rather than an implementation flaw: only a change of architecture gets around it.
The authors show that the condition for stably storing state contradicts the condition under which the gradient does not vanish. Tuning the learning rate or the initialisation therefore cannot fix it. The paper appeared in English in a refereed journal and so influenced the field more than Hochreiter's German-language thesis, though it came three years later.