Back to timeline

Benchmark · November 16, 2022

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

HELM: thirty models, the same forty-two scenarios

On 16 November 2022 Stanford's Center for Research on Foundation Models ran thirty models from twelve organizations through forty-two scenarios, measuring seven things instead of one. Before this, models had on average been evaluated on just 17.9 per cent of those scenarios; after the run, on 96.0 per cent.

Why it matters

Until then comparing two models meant comparing two reports written by different people on different sets, and some prominent models shared not a single scenario. One run under controlled conditions made comparison possible - and put a number on how absent it had been.

The seven metrics are named one by one: accuracy, calibration, robustness, fairness, bias, toxicity and efficiency. They are measured for each of sixteen core scenarios where possible, which came to 87.5 per cent of the time. On top of that, seven targeted evaluations over twenty-six further scenarios: reasoning, disinformation and so on. Forty-two scenarios in all, twenty-one of which had not previously been used in mainstream model evaluation. The number this record exists for: before HELM, models had on average been evaluated on 17.9 per cent of the core scenarios, and some prominent models had none in common. After the run: 96.0 per cent, with all thirty models densely benchmarked on the same scenarios under standardised conditions. The authors name separately what their selection lacks: among other things, question answering for neglected English dialects, and metrics for trustworthiness. Twenty-five top-level findings were published, along with every raw prompt and model completion, so that somebody else's conclusion can be checked against the same data. This record does not claim that HELM made selective reporting impossible; it made it visible, because what had been reported before was now laid out beside it.

Event record

Event date
November 16, 2022
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 21, 2026
Lines
ID
evt-0512

The day of the first version of the preprint, per the arXiv submission history.

Sources

Related events

Records that link to this one