Hugging Face: community-reported evals
On 4 February 2026 Hugging Face launched Community Evals: datasets on the Hub can serve as leaderboards, models keep their own scores, and anyone can submit a result for any model by pull request.
Why it matters
Test results scattered across model cards and papers began to be collected in one place, labelled by who submitted them. An editorial assessment: it exposes discrepancies but does not solve benchmark saturation.
What is said. A dataset registers as a benchmark through an eval.yaml in the Inspect AI format and automatically collects results from across the Hub; MMLU-Pro, GPQA and HLE are already live, starting with a shortlist of four benchmarks. A model's scores live in .eval_results/*.yaml files in its repository; the model author can close a request or hide a result. Any user can submit a result for any model, shown as "community"; badges distinguish author-submitted, community-submitted and independently verified. What the record does not claim. Only the announcement post was read; how many results came in is unknown. The outside report's address led to a different post (30 June 2026 on Every Eval Ever), which is not used.