Transformer
The Transformer used attention without recurrent layers.
Why it matters
It replaced recurrence with attention and made sequence computation more parallelizable.
The paper reports translation and parsing experiments, not a universal capability claim.