vLLM: the memory of attention starts being paged
On 19 June 2023 the first release of vLLM appeared, an engine for serving language models. At its centre is PagedAttention: managing the key-value cache the way operating systems manage paged virtual memory.
Why it matters
Serving a model stopped being bounded by memory held and wasted. The claimed gain is two to four times the throughput at the same latency, which directly changed how many people one model could serve for a given cost.
The paper "Efficient Memory Management for Large Language Model Serving with PagedAttention" was submitted on 12 September 2023 to SOSP 2023, first author Woosuk Kwon, nine authors in all. By its account PagedAttention gives near-zero waste in cache memory and flexible sharing of the cache across requests; the two-to-four-times figure is claimed against FasterTransformer and Orca. The order of events here is the reverse of the usual for a research record: the package became available on 19 June 2023, three months before the paper, and the repository was created earlier still, on 9 February 2023. The record therefore sits on availability rather than publication, because the paper describes something that was already running.