The transformer for images
An image was cut into patches and fed to a transformer as a sequence of words, and given large data it beat convolutional networks.
Why it matters
Two fields that had developed separately for eight years converged on one architecture, the precondition for multimodal models.
A convolution builds assumptions of locality and shift invariance into the network, which helps greatly on small data. The paper showed that given enough volume the model learns these itself, and the built-in assumption becomes a constraint. The consequence runs wider than vision: one architecture for text, images and audio is what made CLIP, DALL-E and everything working across modalities possible.