Helix 2.5: a robot in unfamiliar homes
On 17 September 2026 Figure introduced Helix 2.5, a robot control model pretrained on Index, its dataset of human behaviour, and reported that humanoid robots running it did three household tasks (tidying a living room, folding towels, making a bed) in 30 Bay Area homes where nothing had been collected. Pretraining on Index raised the zero-shot success rate on those tasks from 9% to 56% with the task data held fixed.
Why it matters
A claim that a humanoid can be taught a behaviour once and then sent into homes it has never seen, supported by an ablation that isolates the pretraining. The success criterion is strict: the whole task, no partial credit. The result was obtained and graded by Figure itself, and the company writes that general humanoid robotics is not solved.
What was measured. The failures: 56% success means 44% of trials did not complete the task. The evaluation was blind and the criteria were fixed before the trials; success meant all toys (13-15) in the basket, all towels folded, pillows and comforter at the top of the bed; any human intervention for safety counted as a failure. 'Zero-shot' refers to the evaluation environments and objects: the behaviours themselves were specified with data collected elsewhere. A policy initialised from random weights scored 9% against 56% for the one built on Index, with identical task data, architecture, optimisation and evaluation. Nine per cent is from the page; a news summary seen in a search gave 8%. Other points in the post. Helix 2.5 needed half as much task data as a representative Helix 02 behaviour for the same success, while its coverage grew to 30 unseen homes. A scaling law: four models on nested subsets of Index over an 8x data range; the loss of the largest run was forecast to four decimal places, with a forecast error of 0.54% of the variation. Figure writes that Index accumulates about 35 minutes of new human experience every second and that it has committed $3.5 billion of compute to training Helix (the company's statements). What the record does not claim. The 'first demonstration' claims are the company's, with the caveat 'to our knowledge'. No independent repetition was read; the numbers of homes and tasks are small. The scaling law concerns data volume only, with model size and downstream training held fixed.