Back to timeline

Benchmark · June 22, 2025

The first independent evaluation of generalist policies

RoboArena evaluates robot policies not by a fixed task set but by distributing the work: evaluators at seven institutions freely choose their tasks but must run double-blind pairwise comparisons. Seven policies were compared over more than 600 episodes.

Why it matters

Claims about robot policies had been graded by their own authors, on their own tasks, in their own rooms, so numbers from different laboratories were not comparable. A double-blind protocol across several sites is the first way in this arc to check a manipulation claim from outside the group that made it.

Rather than standardising tasks, environments and locations, the approach distributes evaluation across a network of evaluators who choose what to test on but must compare pairs of policies blind. Preferences from those pairwise comparisons are aggregated into a ranking. The network was instantiated at seven academic institutions on the DROID platform, that is on identical hardware, without which comparing sites would not work. Over more than 600 pairwise real-robot episodes, seven generalist policies were compared, and the paper reports that this ranking is more accurate than conventional centralised evaluation, as well as more scalable and more resilient. The seven institutions are not named in the abstract.

Event record

Event date
June 22, 2025
Timeline date
Event date
Verification
Sources gathered automatically · September 20, 2026
Lines
ID
evt-0452

The date the first version was posted to arXiv.

Sources

Related events