Vision and control trained as one
Sergey Levine, Chelsea Finn, Trevor Darrell and Pieter Abbeel showed that training perception and control jointly, rather than as two separate stages, makes a robot roughly twice as reliable on the hardest contact-rich tasks.
Why it matters
The standard pipeline estimated object pose first and planned motion second, so perception error and control error compounded. Joint training removed that seam and made the visual features answerable to the motion they had to produce.
A network of about 92,000 parameters mapped raw camera images directly to motor torques. Training used a partially observed guided policy search, which turns policy search into supervised learning. Evaluation ran on a PR2 across four tasks requiring close coordination of vision and motion: hanging a coat hanger, fitting a block into a shape-sorting cube, fitting the claw of a toy hammer under a nail, and screwing a cap onto a bottle. On the strictest condition, the visual test, end-to-end training gave 100% on the coat hanger, 87.5% on the shape cube, 78.3% on the hammer and 62.5% on the bottle cap. For comparison, the same architecture using pre-estimated pose as its features gave 83.3%, 40%, 53.3% and 27.5% respectively. On the hardest task, then, the success rate rose from 27.5% to 62.5%.