Back to timeline

Benchmark · April 9, 2024

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

A claimed context length and a real one

On 9 April 2024 NVIDIA published RULER, a set of 13 tasks measuring the length at which a model still actually works. Of ten models tested, only four held 32 thousand tokens, although every one of them claimed that much or more.

Why it matters

Until then long context was checked with needle-in-a-haystack, and nearly everything passed it — in the paper's own table those scores are 100.0 at every length. RULER showed that the test measures nothing, and produced a number that can be set beside the claimed one: the effective length.

To single-needle retrieval the benchmark adds multi-needle variants, multi-hop entity tracing, aggregation across the whole text and question answering — 13 tasks, lengths from 4 to 128 thousand tokens, ten models. The threshold is not arbitrary: it is Llama2-7B's score at 4 thousand tokens, 85.6%. The effective length is the greatest length at which a model still passes it. Claimed against effective: GPT-4, 128K against 64K; Command-R 35B, 128K against 32K; Yi 34B, 200K against 32K; Mixtral 8x7B, 32K against 32K; Mistral 7B, 32K against 16K; ChatGLM 6B, 128K against 4K; LWM 7B, a claimed million against under 4 thousand. The record does not claim the 85.6% threshold is anything natural: the authors call it qualitative and derive it from one reference model. Two non-transformer architectures were also tested, RWKV-v5 and Mamba-2.8B; both degrade markedly by 8 thousand tokens.

Event record

Event date
April 9, 2024
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 21, 2026
Lines
ID
evt-0532

The day the first version of the preprint was submitted. The second followed on 11 April and the third on 6 August 2024; the figures were read in the first.

Sources

Related events