Summary
Figure says its Helix 2.5 neural network performed three whole-body household tasks across 30 previously unseen Bay Area homes without collecting data or adapting in those environments. Index pretraining increased zero-shot success from 9% to 56% in a controlled comparison.
Figure has introduced Helix 2.5, a neural network designed to control a humanoid robot across unfamiliar household environments. In an evaluation spanning 30 previously unseen Bay Area homes, the system performed three long-horizon tasks—tidying living rooms, folding towels and making beds—without collecting data, fine-tuning the model or adapting to the homes and objects used in testing.
The result is a step toward robots that can transfer skills between environments rather than being trained separately for every building. Figure describes it as its first demonstration of zero-shot whole-body generalisation at this scope on a humanoid robot. The claim is based on the company's evaluation and is limited to the tested tasks, homes and setup.
What Helix 2.5 was tested on
Helix 2.5 was pretrained on Index, Figure's dataset of human behaviour, and then adapted to the three household behaviours. The same foundation model supported tasks involving walking, active perception, rigid and deformable objects, bimanual coordination and manipulation.
The evaluation homes were not used for data collection. The objects in the tests were also held out from the task-specification data: this included the toys, towels and bedding used during evaluation. Each task used one fixed checkpoint across all 30 homes, with no model-weight changes based on the evaluation environments or on performance during testing.
The success criteria required completing the full task rather than receiving credit for partial progress. For the living-room task, the robot had to pick up all 13 to 15 toys scattered in the scene and place them in a basket. For towel folding, all towels had to be folded and placed in the basket. For bed making, both pillows and the comforter corners had to be placed at the top of the bed, with the comforter pulled smooth.
These are whole-body tasks because the robot must coordinate several capabilities at once. It may need to walk to locate an object, reposition its stance to reach it, manipulate the object with one or both hands and continually adjust its view of the scene. In a home, furniture and narrow spaces also require the robot to move its body as part of the task rather than operating from a fixed workstation.
Index pretraining was the main performance difference
Figure compared two policies trained with identical task-specific data. One began with random weights, while the other began with the Index-pretrained Helix 2.5 model. The architecture, optimisation settings, downstream data and evaluation process were held fixed.
In blind evaluations, the policy trained from scratch succeeded in 9% of zero-shot trials. The Index-pretrained policy succeeded in 56%. Figure says this sixfold-plus difference indicates that broad pretraining on human behaviour supplied most of the capability needed to transfer the specified tasks to new homes and objects.
The company also reports that Helix 2.5 matched the success rate of a previous Helix 02 policy while using half as much adaptation data. Helix 02 had been trained with data collected directly in the environment where it was evaluated; Helix 2.5 produced the comparison result across unseen homes.
Figure observed qualitative self-correction during long tasks. The robot could step back to reposition itself, change its stance or move around a bed after an unsuccessful action, allowing it to continue working instead of stopping at the first error.
A measured relationship between human data and robot action prediction
Figure trained four models on nested subsets of Index covering an eightfold range of pretraining data. Model size and downstream training were held constant, and the models were assessed using the same held-out robot-action prediction loss.
The loss decreased with each doubling of Index data. Using the smaller training runs, Figure says it predicted the largest run's test loss to four decimal places before that run was trained. The forecasting error was 0.54% of the variation across the full eightfold data range.
This result concerns data scaling with the model size and downstream training fixed. It suggests that larger collections of human behaviour may make transfer to robot actions more predictable, although the reported evaluation covers three behaviours in 30 homes rather than general household work.
Figure says Helix 2.5 is not a solution to general humanoid robotics. Its evidence is concentrated on three defined tasks, a set of 30 Bay Area homes and the company's Index-pretraining approach. The broader significance is that one robot policy could be specified for several behaviours and deployed in new environments without home-specific data collection or fine-tuning.