Back to timeline

Research · February 20, 2025

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

A constant number of cache pages

On 20 February 2025 MIT, Shanghai Jiao Tong University and NVIDIA showed LServe, which applies sparse attention to both prefilling and decoding. Prefilling was accelerated by up to 2.9 times and decoding by 1.3 to 2.1 times over vLLM.

Why it matters

Until then sparse attention was applied to one stage at a time. The point here is not the speedup but the observation underneath it: preserving long-context capability needs only a constant number of cache pages, and that number does not depend on how long the context is.

Half the attention heads are turned into streaming heads, where they cost almost nothing, and this is done at both stages at once. Computation is skipped block-wise, by cache page. The paper's key observation: holding on to long-context capability requires a constant number of KV pages, irrespective of context length. From that follows a hierarchical page selector for each query token, whose selection is reused across tokens, cutting the overhead of selection itself fourfold. Static and dynamic sparsity turned out to be orthogonal, so their speedups multiply rather than replace one another. What the record does not claim. The 2.9 and the 1.3-2.1 are comparisons against vLLM, not against a hardware limit, and the paper states the first as "up to", an upper bound.

Event record

Event date
February 20, 2025
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 22, 2026
Lines
ID
evt-0572

The day the first version of the preprint was submitted. arXiv later added a second; the figures were read in the first.

Sources

Related events

Records that link to this one