One transformer, seven hundred tasks
RT-1 was trained on 130,000 episodes collected over 17 months by a fleet of 13 robots; a single model performed 97% of more than 700 learned instructions and 76% of instructions it had never seen.
Why it matters
The scaling argument became concrete for robots: one model absorbing hundreds of tasks does better than per-task policies, and its generalisation improves with the diversity of the data rather than with engineering for each case.
The work was done by Google and Everyday Robots. The training set was over 130,000 episodes covering more than 700 tasks, collected over 17 months by a fleet of 13 robots. The model performed 97% of the more than 700 instructions it had seen in training, and 76% of those it had not. Robustness was measured separately: 83% on tasks with distractor objects in frame and 59% on tasks with changed backgrounds. Those 700 tasks became, for a long time, the figure against which every later generalist policy was measured. Compared with SayCan, the separate per-skill policies are gone: one model covers them all.