Back to timeline

Research · June 27, 2018

Closed-loop grasping at 96 percent

QT-Opt was trained by reinforcement learning on over 580,000 real grasp attempts; the policy reached 96% success on unseen objects and developed behaviours nobody had specified for it.

Why it matters

Here the grasping strategy stopped being a plan chosen before the motion began and became a policy revised at every frame. Regrasping, probing an object for a better hold and repositioning it before the grasp came out of the optimisation, not out of a designer’s head.

A Google team led by Dmitry Kalashnikov built a scalable self-supervised reinforcement learning framework for vision-based manipulation. A Q-function of over 1.2 million parameters was trained on more than 580,000 real grasp attempts. Perception was RGB only, from a single over-the-shoulder camera. The policy generalised to 96% grasp success on objects absent from training. Besides the rate itself, the paper lists behaviour that emerged on its own: regrasping after a failure, probing an object to find the most effective hold, repositioning an object into an easier pose, other non-prehensile pre-grasp manipulation, and responding to disturbances and perturbations. It continues the 2016 arm farm, replacing supervised prediction with reinforcement learning.

Event record

Event date
June 27, 2018
Timeline date
Event date
Verification
Sources gathered automatically · September 20, 2026
Lines
ID
evt-0430

The date the first version was posted to arXiv.

Sources

Related events