Fourteen arms and 800,000 grasp attempts
Google trained a convolutional network to predict whether a proposed gripper motion would end in a successful grasp, on 800,000 attempts collected over two months by six to fourteen manipulators running in parallel.
Why it matters
Grasping stopped being an engineering problem and became a data problem: the robots labelled their own experience by trying and failing, and the differences between their hardware became a source of robustness rather than a reason to recalibrate.
Sergey Levine, Peter Pastor, Alex Krizhevsky and Deirdre Quillen trained on monocular images only, without camera calibration and without knowledge of the robot pose. That forced the network to observe the spatial relationship between the gripper and the objects in frame for itself. The training material was over 800,000 grasp attempts collected over two months; between 6 and 14 manipulators ran at any one time, differing in camera placement and in hardware. The resulting network was used to servo the gripper in real time, so the robot could correct a grasp mid-motion instead of committing to one plan. Over 100 test attempts with replacement the method gave a 20% failure rate against 43% for an open-loop baseline and 35% for a hand-engineered one. Without replacement, over the first 30 attempts, it gave 17.5% against 33.7% and 50.8%.