Attention for translation
An attention mechanism weighted parts of an input sequence.
Why it matters
It exposed the fixed-vector bottleneck and added soft alignment over source words.
It is an encoder-decoder translation result, not the later Transformer architecture.