Grasping unseen objects, learned from synthetic images
Ashutosh Saxena, Justin Driemeyer, Justin Kearns and Andrew Ng trained an algorithm to find a grasp point directly in an image, using only synthetic pictures for training; on real objects absent from training the robot picked up 87.5% of attempts.
Why it matters
Until then grasping a new object meant building a three-dimensional model of it first. Here there is no model at all: the system predicts a grasp point as a function of the image, and that transferred to objects it had never seen — a transparent glass, a coil of wire, a strangely shaped power horn.
Training used 2,500 synthetic examples from six object classes, generated with computer graphics; not one real image was in the training set. The paper says so outright: "Although training was done on synthetic images, testing was done on the real robotic arm and real objects." The platform was STAIR, Stanford's mobile robot, with a Neuronics Harmonic Arm: 4 kg, 5 degrees of freedom, a parallel plate gripper, positioning accuracy around 1 mm, and a low-quality webcam near the end-effector. Each object was attempted four times. On objects absent from training the average success rate was 87.5%, and 90% on objects similar to the training classes. Mean error in locating the grasp point was 1.85 cm on novel objects and 1.80 cm on similar ones; for comparison the authors cite 1.5 cm for neonate humans. Classifying whether an image patch is the projection of a grasp point scored 94.2% on a synthetic test set. The transparent wine glass was picked up 100% of the time, and equally so when two-thirds full of water. Mugs and jugs, which have to be taken by the handle, scored worst, because the approach trajectory there is narrow.