Figure says its humanoid robot can now carry household skills into homes it has never seen before — but it still gets the job done only a little more than half the time.
The company introduced Helix 2.5 last Thursday and said the Figure 03 robot completed household tasks across 30 previously unseen Bay Area homes without retraining or collecting additional data from those locations.
Across 420 trials involving bed making, towel folding, and toy tidying, Figure reported 237 successful completions, or 56%. The result shows progress in generalization, but it also highlights the reliability gap that remains before household robots can be trusted to work unsupervised.
Figure calls the test zero-shot because the homes, layouts, and objects were unseen during evaluation. The chores themselves were not new; Helix 2.5 had already learned those task categories from training data collected elsewhere.
Index appears to be driving the generalization gains
The biggest technical claim is that Figure’s Index dataset of human behavior can give robots generalizable physical behavior patterns that transfer between environments.
In a controlled comparison, Figure trained two policies with the same architecture, task data, optimization, and evaluation. The difference was Index pretraining. The model trained from scratch succeeded in 9% of trials, compared with 56% for the Index-pretrained model.
That suggests the improvement was not simply the result of giving Helix 2.5 more examples of the three chores. Figure says no single evaluation task represents more than 1.90% of Index’s pretraining data. The company also says Helix 2.5 required half as much adaptation data as a comparable Helix 02 behavior while achieving similar performance and extending that behavior to 30 unfamiliar homes.
The more important number may be 44%
The 56% completion rate shows progress, but the inverse number may matter more for anyone imagining a robot working unsupervised at home: Helix 2.5 failed 44% of its trials under Figure’s all-or-nothing scoring criteria.
That does not mean every failure was equally serious, but it does show that generalization is not the same as dependable everyday performance.
Toy tidying was particularly difficult, succeeding in only 40% of attempts. A robot that fails nearly half of its complete tasks may still demonstrate useful capabilities, but reliability becomes a major issue when the job is happening in someone’s home without an operator. Figure itself cautions that general humanoid robotics is not solved.
Why the data scaling matters
Figure reported another potentially important result: increasing Index pretraining data across four runs produced predictable improvements in downstream robot-action prediction. The experiments covered an eightfold range of data, and Figure says smaller runs predicted the largest run’s test loss to four decimal places.
That does not prove that doubling the data will double household-task success. The scaling experiment measured action-prediction loss, not real-world completion rates. Still, it gives Figure a measurable framework for deciding whether additional human data and computing resources are improving the underlying model.
What this means for users
For consumers, Helix 2.5 is less about buying a home robot today and more about whether robots can eventually avoid being trained separately for every household.
If Figure’s approach works at much higher reliability, a future robot could potentially learn a chore once and carry that skill into different homes, furniture layouts, and objects. That could reduce the amount of setup required before deployment.
The immediate limitation is reliability. The 30-home test demonstrates generalization, but the reported 56% completion rate leaves a large gap between showing that a robot can handle unfamiliar environments and proving that it can reliably work there every day.
According to Figure, Index generates approximately 35 minutes of new human experience per second, and $3.5 billion in compute has been allocated to Helix’s training. The key uncertainty remains whether expanding data and computing power can transform current partial generalization into reliable, automated household systems.
The remaining hurdle is reliability. Figure has shown that Helix 2.5 can generalize beyond familiar environments; the harder challenge is turning that 56% completion rate into performance people can trust every day.
Other news: AI avatar Tilly Norwood unexpectedly switched from English to Cantonese during an interview, raising questions about how companies monitor, audit, and maintain human control over public-facing AI agents.
Read the full article here