Back to timeline

Research · July 28, 2023

A web-trained model outputs robot actions

RT-2 is a vision-language-action model: a vision-language model pre-trained on the web was fine-tuned so that it emits robot actions as text tokens. On unseen objects and environments, success rose from RT-1’s 32% to 62% across more than 6,000 trials.

Why it matters

Knowledge that had existed only in text and images — what an object is and which objects go together — reached the output the robot moves by, without anyone teaching those concepts on a robot. The model could act on an instruction about a thing it had never held.

Actions are represented as text tokens, so a vision-language model can emit them by the same mechanism it emits words. On scenarios with unseen objects, backgrounds and environments, success rose from RT-1’s 32% to 62%. Evaluation covered more than 6,000 robotic trials. On what the authors called emergent skills — symbol understanding, reasoning and human recognition — the improvement over baselines is roughly threefold. On the Language-Table simulation benchmark the model reached 90% against a previous best of 77%. The blog and the project page state the generalisation differently: the blog reports more than a threefold improvement, the page about twofold on general generalisation and threefold specifically on emergent skills; these are different slices of the evaluation.

Event record

Event date
July 28, 2023
Timeline date
Event date
Verification
Sources gathered automatically · September 20, 2026
Lines
ID
evt-0439

The arXiv first version and the Google DeepMind post are both dated that day.

Sources

Related events

Records that link to this one