<p>Figure introduced Helix 2.5 on September 17, 2026, with a result aimed at one of humanoid robotics’ hardest practical questions: can a robot carry a learned behavior into unfamiliar homes and objects? The company reports an evaluation across 30 real homes in the San Francisco Bay Area, but the evidence describes a controlled test rather than routine household deployment.</p><h2>Three behaviors, one model</h2><p>Helix 2.5 was pretrained on Index, Figure’s global-scale dataset of human behaviors. From that single foundation model, Figure produced behaviors for tidying living rooms, folding towels, and making beds. Those tasks combine perception, locomotion, manipulation, bimanual coordination, and whole-body control, so the result concerns the coordination of a humanoid system rather than an isolated picking demonstration.</p><p>The evaluation used real homes and objects the robot had not encountered before. Figure says it collected no data in the evaluation homes and performed no fine-tuning or adaptation in those environments or on the objects being handled. The reported result is therefore about transfer: the robot was asked to apply learned behaviors without tailoring the model to each location.</p><h2>What the evaluation tested</h2><p>Figure used one fixed checkpoint in all 30 homes. It was not selected using trajectory data or evaluation-performance data. The company also says that none of the evaluation toys, towels, or bedding appeared in the task-specification data, based on an AI-model check followed by human review. These conditions make the claim more specific than a video showing a robot succeeding once in a prepared room.</p><h2>A controlled comparison</h2><p>In a controlled comparison, the Index-pretrained policy reached a 56% zero-shot success rate, versus 9% for a policy trained from scratch. Figure kept the task data, architecture, optimization, hyperparameters, downstream training, and evaluation fixed. Helix 2.5 also matched the success rate of a comparable Helix 02 policy while using half as much adaptation data and generalizing zero-shot to the 30 new homes.</p><p>For researchers and companies assessing humanoid systems, the useful result is a measurable demonstration of task and environment transfer. It does not establish a general household capability: the evaluation covers three behaviors under one setup, and the reported success rate is not a safety validation for home operation.</p><h2>Where the claim stops</h2><p>Figure itself says the result does not mean general humanoid robotics is solved. That caution fits the evidence. Helix 2.5 has shown zero-shot generalization across a defined set of household tasks and 30 unfamiliar homes; on these facts alone, it has not established dependable operation across every home, object, or task.</p><p>Helix 2.5 is best read as a concrete evaluation milestone, not proof that humanoid robots are ready for unsupervised domestic work. Its importance lies in the protocol: one pretrained model, one fixed checkpoint, unfamiliar environments, and a controlled baseline make the improvement easier to interpret than a forward-looking promise.</p><section class="media-fleet-sources"><h2>Official sources</h2><ul><li><a href="https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization">Official source: figure.ai</a></li></ul></section><aside class="media-fleet-related"><h2>Related reading</h2><ul><li><a href="https://rentbuyrobot.com/article/spot-5-2-connects-industrial-inspections-to-ai-agents">Spot 5 2 Connects Industrial Inspections To Ai Agents</a></li><li><a href="https://rentbuyrobot.com/article/universal-robots-welding-cobots-beyond-cart-fabtech-2024">Universal Robots Welding Cobots Beyond Cart Fabtech 2024</a></li></ul></aside>
PrototypeReport
Figure reports Helix 2.5 transferred three household behaviors to 30 unfamiliar homes
Figure says Helix 2.5 achieved a 56% zero-shot success rate across three household tasks in 30 previously unseen San Francisco Bay Area homes, presenting an evaluation result rather than a domestic deployment claim.

