Bimanual teleoperation for twenty thousand dollars
ALOHA is a two-armed teleoperation rig built within a stated budget of 20,000 dollars; from demonstrations collected on it the ACT algorithm learned six fine manipulation tasks from 10 minutes of demonstration each, at 80-90% success.
Why it matters
The bottleneck in imitation learning was who could afford to collect demonstrations at all. A rig at 20,000 rather than hundreds of thousands moved that from industrial laboratories to ordinary university groups, and 10 minutes per task changed the unit of effort for a new skill.
Tony Z. Zhao, Vikash Kumar, Sergey Levine and Chelsea Finn built the rig and wrote Action Chunking with Transformers to learn from demonstrations collected on it. Six real-world tasks: opening a translucent condiment cup, slotting a battery, threading velcro, prepping tape, putting on a shoe and sliding a ziploc bag. Success was 80-90%; the project page gives individual tasks at 96%, 84%, 64% and 92%. The 20,000 dollar budget is named on the project page rather than in the abstract; the abstract carries the tasks, the demonstration time and the success range.