A web-trained model outputs robot actions
RT-2 is a vision-language-action model: a vision-language model pre-trained on the web was fine-tuned so that it emits robot actions as text tokens. On unseen objects and environments, success rose from RT-1’s 32% to 62% across more than 6,000 trials.
Why it matters
Knowledge that had existed only in text and images — what an object is and which objects go together — reached the output the robot moves by, without anyone teaching those concepts on a robot. The model could act on an instruction about a thing it had never held.
Actions are represented as text tokens, so a vision-language model can emit them by the same mechanism it emits words. On scenarios with unseen objects, backgrounds and environments, success rose from RT-1’s 32% to 62%. Evaluation covered more than 6,000 robotic trials. On what the authors called emergent skills — symbol understanding, reasoning and human recognition — the improvement over baselines is roughly threefold. On the Language-Table simulation benchmark the model reached 90% against a previous best of 77%. The blog and the project page state the generalisation differently: the blog reports more than a threefold improvement, the page about twofold on general generalisation and threefold specifically on emergent skills; these are different slices of the evaluation.