NVIDIA H100
The new generation got a dedicated transformer unit: hardware designed for one specific model architecture.
Why it matters
The circle closed: first the method was fitted to the hardware, now hardware is designed for a method named in a 2017 paper.
The unit picks precision automatically for each transformer layer, speeding training without losing quality. The moment is telling: the name of a neural network architecture appears in a chip's specification. Demand for these accelerators in late 2022 exceeded supply so far that access to them became the limiting factor for whole companies.