Compressive memory instead of a growing KV cache
On 10 April 2024 Google described Infini-attention: local attention and a long-term compressive memory inside one transformer block. A 1-billion-parameter model fine-tuned on 5-thousand-token segments retrieved a key hidden in a sequence of a million tokens.
Why it matters
Until then a longer context meant a proportionally larger KV cache, so memory grew with length and had to hit a wall somewhere. Here the older part of the sequence is folded into a fixed-size matrix, and length stops being a question of memory.
The mechanism adds no separate weight matrices: it reuses the Q, K and V vectors of the same attention layer — the query reads from compressive memory, the key and value update it. The segment length in the experiments is 2 thousand tokens. The figure 114 belongs to one specific comparison: Infini-Transformer has 114 times fewer memory parameters (1.6 million against 183 million) than a Memorizing Transformer with a vector-retrieval KV store of length 65K at its ninth layer, while giving better perplexity. Tested on PG19 and Arxiv-math, with 1B and 8B models. Limits. The perfect passkey retrieval at a million tokens in Table 3 belongs to the Linear + Delta variant; plain Linear drops to 96/94/100 at a million depending on where in the text the key is hidden. Book summarisation at 500 thousand tokens (BookSum) is summarisation, not reasoning across the whole length; the record does not claim the model holds an arbitrary task at a million tokens.