Back to timeline

Research · June 12, 2017

Transformer

The Transformer used attention without recurrent layers.

Why it matters

It replaced recurrence with attention and made sequence computation more parallelizable.

The paper reports translation and parsing experiments, not a universal capability claim.

Event record

Event date
June 12, 2017
Timeline date
Event date
Verification
Sources gathered automatically · September 11, 2026
Lines
ID
evt-0011

Sources

Related events

Records that link to this one