Closed-loop grasping at 96 percent
QT-Opt was trained by reinforcement learning on over 580,000 real grasp attempts; the policy reached 96% success on unseen objects and developed behaviours nobody had specified for it.
Why it matters
Here the grasping strategy stopped being a plan chosen before the motion began and became a policy revised at every frame. Regrasping, probing an object for a better hold and repositioning it before the grasp came out of the optimisation, not out of a designer’s head.
A Google team led by Dmitry Kalashnikov built a scalable self-supervised reinforcement learning framework for vision-based manipulation. A Q-function of over 1.2 million parameters was trained on more than 580,000 real grasp attempts. Perception was RGB only, from a single over-the-shoulder camera. The policy generalised to 96% grasp success on objects absent from training. Besides the rate itself, the paper lists behaviour that emerged on its own: regrasping after a failure, probing an object to find the most effective hold, repositioning an object into an easier pose, other non-prehensile pre-grasp manipulation, and responding to disturbances and perturbations. It continues the 2016 arm farm, replacing supervised prediction with reinforcement learning.