Back to timeline

Research · June 17, 2021

Low-rank adaptation (LoRA)

On 17 June 2021 Edward Hu and seven co-authors at Microsoft posted LoRA: the pre-trained weights are frozen and trainable rank decomposition matrices are injected into each layer of the Transformer. For GPT-3 175B the first version reports 10,000 times fewer trainable parameters and a three times lower computation hardware requirement than full fine-tuning, at quality on par or better, and a checkpoint that shrinks from 350 GB to 35 MB.

Why it matters

Adapting a very large model to a task stopped meaning a full copy of it per task: the adapted part weighs megabytes and is swapped on a machine that holds the frozen weights, and because the low-rank product can be merged back into the weights, inference is no slower. QLoRA in 2023 is LoRA applied to a model quantised to 4 bits.

In the body of the first version, for GPT-3: VRAM use during training falls from 1.2 TB to 350 GB, and at rank r = 4 the checkpoint falls roughly 10,000-fold, from 350 GB to 35 MB; training is 25 per cent faster. The abstract of the first version says 'the computation hardware requirement by 3 times'; the second version of 16 October 2021 changes this to 'the GPU memory requirement by 3 times'. The record uses the first wording.

Event record

Event date
June 17, 2021
Timeline date
Event date
Verification
Sources gathered automatically · September 24, 2026
Lines
ID
evt-0670

The first of two versions of arXiv:2106.09685, 17 June 2021; the second is of 16 October 2021. Wording and figures are from the first.

Sources

Related events