Back to timeline

Research · May 23, 2023

QLoRA

The method made it possible to fine-tune a 65-billion-parameter model on one graphics card, holding the weights in four bits and changing only small additional matrices.

Why it matters

Fine-tuning a large model stopped requiring a cluster, and open models became workable for anyone with a gaming card.

Two ideas together: quantising the base weights to four bits and training only low-rank additions. The base model itself does not change, so memory is spent on it once. The quality loss proved slight. The consequence is ecological: the community of fine-tuned open-model variants that grew after LLaMA became technically possible.

Event record

Event date
May 23, 2023
Timeline date
Event date
Verification
Sources gathered automatically · September 17, 2026
Lines
ID
evt-0276

Preprint of 23 May 2023.

Sources

Related events

Records that link to this one