Low-rank adaptation (LoRA)
On 17 June 2021 Edward Hu and seven co-authors at Microsoft posted LoRA: the pre-trained weights are frozen and trainable rank decomposition matrices are injected into each layer of the Transformer. For GPT-3 175B the first version reports 10,000 times fewer trainable parameters and a three times lower computation hardware requirement than full fine-tuning, at quality on par or better, and a checkpoint that shrinks from 350 GB to 35 MB.
Why it matters
Adapting a very large model to a task stopped meaning a full copy of it per task: the adapted part weighs megabytes and is swapped on a machine that holds the frozen weights, and because the low-rank product can be merged back into the weights, inference is no slower. QLoRA in 2023 is LoRA applied to a model quantised to 4 bits.
In the body of the first version, for GPT-3: VRAM use during training falls from 1.2 TB to 350 GB, and at rank r = 4 the checkpoint falls roughly 10,000-fold, from 350 GB to 35 MB; training is 25 per cent faster. The abstract of the first version says 'the computation hardware requirement by 3 times'; the second version of 16 October 2021 changes this to 'the GPU memory requirement by 3 times'. The record uses the first wording.