The attention cache at two bits
On 25 January 2025 Tsinghua and Huawei showed RotateKV, which compresses the key-value cache to two bits by rotating the activation space. Peak cache memory fell by a factor of 3.97, and perplexity on WikiText-2 with LLaMA-2-13B worsened by less than 0.3.
Why it matters
As context length and batch size grow, the bottleneck stops being the size of the model and becomes the attention cache. The work showed the cache can be held at two bits without visible loss — so the same model on the same hardware can serve a queue five times longer.
The device rests on rotating the activation space, with two refinements: the rotations are adapted to the particular channels carrying outliers by reordering them, and the so-called attention sinks are protected separately, identified by their massive activations. Figures from the first version: under 0.3 points of perplexity degradation on WikiText-2 with LLaMA-2-13B at two bits; under 1.7 per cent lost on GSM8K; a 3.97-fold reduction in peak cache memory; support for 5.75 times larger batches; a 2.32-fold speedup in decoding. What the record does not claim. The measurements are on LLaMA-2 family models and the named benchmarks, not on an arbitrary model. The compression is of the cache, not the weights: different parts of memory and different techniques.