LSTM
LSTM was proposed for sequence dependencies.
Why it matters
LSTM addressed the difficulty of retaining information over long sequences.
Its result concerns a recurrent architecture, not every later language model.
Research · November 1997
Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.
LSTM was proposed for sequence dependencies.
LSTM addressed the difficulty of retaining information over long sequences.
Its result concerns a recurrent architecture, not every later language model.
Neural Computation · Published November 1997
A recurrent architecture is trained by the gradient method popularised earlier.
The authors explicitly use multilayer LSTM networks to encode the input sequence and generate a translation.
Sequence to Sequence LearningThe recurrent scheme whose limitation LSTM later solves.
LSTM in 1997 is the direct answer to this diagnosis.
With Hochreiter's work it justifies the need for a new recurrent scheme.
The model belongs to the line of recurrent schemes with a controlled state.
The preprint says the new hidden unit was motivated by the LSTM unit but is much simpler.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation