Figure says two humanoid robots equipped with its Helix-02 system reset a bedroom in under two minutes using a single learned Vision-Language-Action policy. In the company’s May 8, 2026 announcement, the robots are shown opening a door, hanging clothes, putting away headphones, closing a book, taking out trash, moving a chair and working together to make a bed.

The notable claim is not only that two robots performed several household tasks. Figure says they coordinated using their own camera observations, without a shared planner, explicit message passing or a central coordinator. As Figure describes it, the demonstration used neither scripted handoffs nor separate task-specific controllers for the sequence, rather than establishing that no scripted components of any kind were used.

Figure does not report the number of trials, failure rate, repeatability, safety incidents, training-data scale or the extent of human supervision.

Contents

What Figure demonstrated

Figure describes the sequence as a bedroom reset performed by two Helix-02-equipped humanoids. The reported actions include:

  • opening a door using whole-body coordination;
  • walking between different locations in the room;
  • hanging clothing on a coat tree;
  • placing headphones on a stand;
  • pushing an office chair beneath a desk;
  • closing a book;
  • operating a trash can’s foot pedal;
  • manipulating bedding while the two robots work around the same bed.

The sequence combines locomotion, balance, sensing and manipulation of different object types. Some objects, such as the chair and book, are relatively rigid. Clothing and bedding are deformable: their shape and contact points change as the robots handle them.

Figure says the robots completed the sequence without scripted handoffs between subtasks or task-specific controllers. It also says the same underlying approach has previously been used for logistics tasks, laundry folding, kitchen cleanup and living-room tidying, with additional data rather than a change to the core algorithm.

The company describes this as, to its knowledge, the first example of a single learned neural network performing multi-humanoid collaborative “locomanipulation” directly from pixels to actions.

How the coordination is described

A Vision-Language-Action system generally connects visual perception and task-level semantic or language information to physical robot actions. In this case, Figure says both robots run one learned policy. The announcement does not disclose the policy’s exact neural-network architecture, language interface, camera configuration, training data or control-loop implementation.

The coordination mechanism Figure describes is different from a conventional centrally planned multi-robot system. Rather than using a central system to assign actions or exchanging explicit messages, each robot observes the room through its own cameras and infers the other robot’s intent from its movements.

That suggests a decentralised form of coordination based on observation. Each robot still has to repeatedly decide what to do as the room changes and as the other robot moves. This is more difficult than running two independent tasks in separate parts of a building: the robots must avoid interfering with one another and adapt when their actions change the shared scene.

The announcement also says the robots did not use scripted handoffs between subtasks. That matters because a handoff could otherwise specify in advance which robot should act next or how one robot should respond to a known intermediate state. Figure’s description instead attributes the sequence to a common learned policy operating across the changing environment.

Why the task is technically demanding

Whole-body manipulation is different from controlling a stationary robotic arm. It combines arm and hand movements with foot placement, balance and body posture. Opening a door, for example, may require a robot to position its body, maintain balance while applying force and coordinate its arm with its movement.

Figure’s sequence also includes dynamic one-leg balance, rigid-object manipulation and deformable-object manipulation. Bedding is particularly challenging because it does not retain a fixed geometric shape. As a comforter is lifted, pulled or placed, its configuration changes, making it harder to predict where the next useful contact point will be.

The shared bed adds a multi-robot coordination problem. Two robots must act around the same object without blocking one another or producing conflicting forces. The room also contains furniture and objects distributed across different locations, so the demonstration combines movement through the environment with dexterous actions rather than testing a single isolated skill.

This is why the reported result is relevant to general-purpose humanoid research. A household robot would need to combine perception, navigation, balance and manipulation while operating in spaces designed for people. However, a successful demonstration of that combination does not by itself show that the policy will generalise reliably to other rooms, object arrangements or types of bedding.

What the demonstration does not establish

The evidence is a company announcement and demonstration, not a peer-reviewed research publication or independent benchmark. Figure does not provide a formal comparator, sample size, confidence estimates or a task-completion distribution.

Several important details remain unknown:

  • how many bedroom-reset attempts were conducted;
  • whether the reported run was continuous and unedited;
  • whether any resets, interventions or setup occurred outside the shown sequence;
  • how often individual tasks failed;
  • how variable the completion time was;
  • how accurately objects were placed;
  • how the system recovered from errors;
  • how much human supervision or remote assistance was involved;
  • whether the environment required preparation for the demonstration.

The “under two minutes” figure should therefore be read as a reported demonstration result, not as an established average operating time.

The announcement also does not establish safe operation around people, children or pets. It provides no evidence about performance under changing lighting, unfamiliar furniture layouts, different objects or unpredictable human activity. Nor does it establish commercial availability, pricing, regulatory approval or readiness for unsupervised household use.

What to watch next

The most useful follow-up evidence would be repeatability data across different bedrooms, layouts, furniture arrangements and bedding types. Task-completion rates, failure cases, timing variability and recovery behaviour would show whether the capability extends beyond a single demonstration.

Further technical disclosure would also clarify what is attributable to Helix-02 hardware and what comes from the learned policy. Relevant details include the policy architecture, training-data scale, sensor configuration, control frequency, safety mechanisms and the role of demonstrations or human supervision during training and deployment.

Tests with people or other robots moving unpredictably nearby would be important for assessing real shared-space operation. Until such evidence is available, the Helix-02 bedroom sequence is best understood as a technically ambitious company demonstration of multi-humanoid coordination, rather than proof of a reliable household product.

Sources