A language model plans a robot’s actions
SayCan paired a language model’s proposal of what to do with a learned value function’s estimate of what the robot can actually do right now; across 101 tasks given in natural language this gave 84% successful plans and 74% successful executions.
Why it matters
A language model was joined to a robot without pretending it knew anything about the body: the value function vetoed steps the robot could not perform. That division — planner above, feasibility check below — is what the vision-language-action models later collapsed into a single network.
The work was done by Google and Everyday Robots and run on a mobile manipulator in an office kitchen. Evaluation covered 101 tasks specified by natural-language instructions. The figures of 84% for planning and 74% for execution were published with the PaLM-SayCan update of 16 August 2022; the project page states these are 14 and 13 points higher than the first version, posted on 4 April 2022. The April release therefore scored about that much lower, and the August numbers must not be attached to the April date. The method ranks candidate next steps twice: the language model by usefulness towards the instruction, the value function by attainability in the current state, and the step chosen is the one that scores well on both.