Back to timeline

Announcement · February 4, 2026

Hugging Face: community-reported evals

On 4 February 2026 Hugging Face launched Community Evals: datasets on the Hub can serve as leaderboards, models keep their own scores, and anyone can submit a result for any model by pull request.

Why it matters

Test results scattered across model cards and papers began to be collected in one place, labelled by who submitted them. An editorial assessment: it exposes discrepancies but does not solve benchmark saturation.

What is said. A dataset registers as a benchmark through an eval.yaml in the Inspect AI format and automatically collects results from across the Hub; MMLU-Pro, GPQA and HLE are already live, starting with a shortlist of four benchmarks. A model's scores live in .eval_results/*.yaml files in its repository; the model author can close a request or hide a result. Any user can submit a result for any model, shown as "community"; badges distinguish author-submitted, community-submitted and independently verified. What the record does not claim. Only the announcement post was read; how many results came in is unknown. The outside report's address led to a different post (30 June 2026 on Every Eval Ever), which is not used.

Event record

Event date
February 4, 2026
Timeline date
Event date
Verification
Sources gathered automatically · September 29, 2026
Lines
ID
evt-0953

Sources

Related events

Antecedents for this event are still being researched.