Back to timeline

Research · October 2020

The transformer for images

An image was cut into patches and fed to a transformer as a sequence of words, and given large data it beat convolutional networks.

Why it matters

Two fields that had developed separately for eight years converged on one architecture, the precondition for multimodal models.

A convolution builds assumptions of locality and shift invariance into the network, which helps greatly on small data. The paper showed that given enough volume the model learns these itself, and the built-in assumption becomes a constraint. The consequence runs wider than vision: one architecture for text, images and audio is what made CLIP, DALL-E and everything working across modalities possible.

Event record

Event date
October 2020
Timeline date
Event date
Verification
Sources gathered automatically · September 17, 2026
Lines
ID
evt-0241

Preprint of 22 October 2020; presented at ICLR 2021.

Sources

Related events

Records that link to this one