A claimed context length and a real one
On 9 April 2024 NVIDIA published RULER, a set of 13 tasks measuring the length at which a model still actually works. Of ten models tested, only four held 32 thousand tokens, although every one of them claimed that much or more.
Why it matters
Until then long context was checked with needle-in-a-haystack, and nearly everything passed it — in the paper's own table those scores are 100.0 at every length. RULER showed that the test measures nothing, and produced a number that can be set beside the claimed one: the effective length.
To single-needle retrieval the benchmark adds multi-needle variants, multi-hop entity tracing, aggregation across the whole text and question answering — 13 tasks, lengths from 4 to 128 thousand tokens, ten models. The threshold is not arbitrary: it is Llama2-7B's score at 4 thousand tokens, 85.6%. The effective length is the greatest length at which a model still passes it. Claimed against effective: GPT-4, 128K against 64K; Command-R 35B, 128K against 32K; Yi 34B, 200K against 32K; Mixtral 8x7B, 32K against 32K; Mistral 7B, 32K against 16K; ChatGLM 6B, 128K against 4K; LWM 7B, a claimed million against under 4 thousand. The record does not claim the 85.6% threshold is anything natural: the authors call it qualitative and derive it from one reference model. Two non-transformer architectures were also tested, RWKV-v5 and Mamba-2.8B; both degrade markedly by 8 thousand tokens.