Open weights for a vision-language-action model
OpenVLA, at 7 billion parameters, was released with open weights; across 29 tasks it beat the closed 55-billion-parameter RT-2-X by 16.5 percentage points of absolute success while using roughly seven times fewer parameters.
Why it matters
Until then every frontier robot policy sat inside one company. A model anyone could download that also beat the closed state of the art on a shared task set meant a laboratory without a robot fleet could start from the frontier rather than from zero.
The work was done jointly by Stanford, Berkeley, Toyota Research Institute and Google DeepMind. The model is built on a Llama 2 language model with a visual encoder fusing DINOv2 and SigLIP features, and trained on 970,000 real-world robot demonstrations. Evaluation covered 29 tasks and several different robot embodiments. The margin over RT-2-X is 16.5 percentage points of absolute success rate, while RT-2-X has 55 billion parameters against 7 billion here. The model can be fine-tuned on consumer hardware, which is what makes the openness practical rather than nominal.