Grasping learned from synthetic data
Dex-Net 2.0 was trained on 6.7 million synthetic point clouds generated from 3D models, with no attempts on a real robot at all — and the resulting policy gave 99% precision on forty unseen household objects.
Why it matters
It turned out the data for grasping need not be collected by robots at all: physics and geometry can generate it, and the policy still transfers to hardware and to novel objects. That decoupled progress in grasping from the cost of robot-hours.
Jeffrey Mahler and colleagues at Berkeley, with Juan Aparicio Ojea of ABB and Ken Goldberg, trained a grasp-quality convolutional network on a synthetic set of 6.7 million point clouds, grasps and analytic grasp metrics generated from thousands of 3D models. No real-robot data entered training. The policy was then tested on an ABB YuMi in over 1,000 physical trials. On eight known objects it gave 93% success. On a set of 40 novel household objects it gave 99% precision: one false positive out of 69 grasps classified as robust. This is the opposite approach to the arm farm, which attacked the same problem with the sheer number of real attempts where this one used simulation.