Diffusion policies for visuomotor control
Instead of predicting a single action, Diffusion Policy represents the robot’s motion as a diffusion process over action sequences; across 12 tasks from four existing benchmarks this gave an average improvement of 46.9% over prior methods.
Why it matters
Behaviour cloning was undone by averaging: when two motions are correct, a regression model returns something between them, which is a third and wrong one. Modelling the distribution of actions rather than its mean removed that fault, and diffusion became the standard action output for the generalist policies that followed.
Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake and Shuran Song worked together across Columbia University, Toyota Research Institute and MIT. The policy is represented as a conditional denoising diffusion process over action sequences, which allows it to handle multimodal distributions where several different motions are equally correct. Evaluation ran on 12 different tasks from 4 different manipulation benchmarks, that is against other people’s already published baselines rather than against an earlier version of itself. The average improvement was 46.9%.