Figure has announced Index, a company-exclusive system for collecting and processing videos of people performing real-world tasks for robot training. The company says the pipeline is intended to supply data for Helix, its robotics artificial-intelligence system.

In an announcement dated 25 August 2026, Figure said Index had already been operating through an app launched in stealth four months earlier. The app is now being launched publicly on Google Play and the App Store. Figure says contributors, whom it calls “Creators”, record activities including cooking, cleaning, laundry, logistics-centre work, restaurant work, factory work and office work.

The announcement reports more than 264,000 downloads across 108 countries, more than 44,000 weekly active users and over 16 million uploaded videos.

Contents

What Figure is launching

Index is designed to collect human-generated video of physical activities, objects and environments rather than relying only on data recorded by robots. In machine learning, training data is used to adjust a model’s parameters—the numerical values that influence how it responds to inputs. Greater variation in the examples can help a system deal with conditions that differ from its training data, although the result depends on data quality, coverage and the training method.

For a robotics system, video can provide visual and temporal information: what objects are present, how they move and how a task unfolds over time. A recording of a person folding laundry, for example, may show a sequence of actions and changes in object position. It does not by itself establish that a robot can reproduce the task safely or reliably.

Figure says Index is intended to address what it describes as a data problem in robotics: obtaining large quantities of varied physical-world examples for training. The company’s announcement presents the system as Figure-exclusive and links it to future training of Helix.

The supplied announcement does not specify the total number of hours collected. It also does not report robot trials, task-success rates, error rates or commercial deployment results based on Index data.

How Index processes the recordings

Figure describes a five-stage processing pipeline:

  1. filtering;
  2. fraud review;
  3. deduplication;
  4. rebalancing; and
  5. annotation.

The company says automated filters first screen submissions for technical, visual and semantic quality. Human analysts then audit samples at the user level to identify deliberate attempts to evade those filters.

For deduplication, Figure says it converts video segments into embeddings. An embedding is a numerical representation used to compare items by their learned similarity. Segments that exceed a similarity threshold against previously accepted data are discarded, according to the company, with the aim of preserving variety rather than repeatedly adding near-identical recordings.

Figure says it then rebalances the accepted material using task quotas and embedding-based clusters. Task quotas can control how much of the collection is devoted to particular activities, while clusters group data according to similarity in the representation space. The announcement does not define the quotas, the clustering method or the threshold used for rejecting similar segments.

The final stated stage is annotation. Figure says it generates hierarchical text captions for every episode. These captions are intended to describe the content at multiple levels, although the supplied announcement does not provide examples of the captions or explain whether they are generated automatically, checked by people, or both.

What the reported scale means

Figure reports that the app is processing 30 minutes of uploaded video every second. In the company’s description, that is equivalent to 4.9 years of human work being uploaded each day.

The company also reports that it has paid Creators $15 million to date. It gives the following diversity figures for every 1,000 hours collected:

  • 373 unique tasks;
  • 1,146 unique manipulated objects; and
  • 116 unique environments.

These measures suggest that Figure is tracking more than the number of videos. The company is also attempting to quantify how many different activities, objects and settings appear in the collection. However, the announcement does not define the counting rules or explain how duplicates and near-duplicates are treated beyond the stated deduplication process.

The number of environments is not necessarily a measure of global representativeness. The announcement does not disclose the geographic distribution or demographic composition of contributors, nor how closely the recordings reflect the range of homes, workplaces, tools and physical conditions in which robots might eventually operate.

The announcement describes Index in terms such as the largest and most diverse physical dataset, but it does not provide an independently verified comparison with other robot-training datasets. The reported downloads, active users, uploads, processing rate, payments and diversity measures are all company claims.

Planned expansion and robotics-as-a-service

Figure says it plans to scale its data-collection efforts by 100 times and is committed to spending more than $1 billion on data and compute over the following 12 months. Both statements are future commitments, not completed results.

The company also says Index is intended to lay the groundwork for a future robots-as-a-service model, in which customers would order robotic capability as a service rather than necessarily buying and operating the machines themselves.

Whether that model becomes practical will depend on more than the volume of collected video. Figure would need to demonstrate that the data can support reliable physical action, address safety and privacy requirements, and produce measurable improvements in real-world robot performance. The current announcement does not yet provide that evidence.

Sources