The first independent evaluation of generalist policies
RoboArena evaluates robot policies not by a fixed task set but by distributing the work: evaluators at seven institutions freely choose their tasks but must run double-blind pairwise comparisons. Seven policies were compared over more than 600 episodes.
Why it matters
Claims about robot policies had been graded by their own authors, on their own tasks, in their own rooms, so numbers from different laboratories were not comparable. A double-blind protocol across several sites is the first way in this arc to check a manipulation claim from outside the group that made it.
Rather than standardising tasks, environments and locations, the approach distributes evaluation across a network of evaluators who choose what to test on but must compare pairs of policies blind. Preferences from those pairwise comparisons are aggregated into a ranking. The network was instantiated at seven academic institutions on the DROID platform, that is on identical hardware, without which comparing sites would not work. Over more than 600 pairwise real-robot episodes, seven generalist policies were compared, and the paper reports that this ranking is more accurate than conventional centralised evaluation, as well as more scalable and more resilient. The seven institutions are not named in the abstract.