QLoRA
The method made it possible to fine-tune a 65-billion-parameter model on one graphics card, holding the weights in four bits and changing only small additional matrices.
Why it matters
Fine-tuning a large model stopped requiring a cluster, and open models became workable for anyone with a gaming card.
Two ideas together: quantising the base weights to four bits and training only low-rank additions. The base model itself does not change, so memory is spent on it once. The quality loss proved slight. The consequence is ecological: the community of fine-tuned open-model variants that grew after LLaMA became technically possible.