HumanEgo self-reported: 30 minutes of human video per task trains manipulation policies at 92.5% success
HumanEgo converts egocentric human video into entity-level hand-object representations to train flow-matching policies, self-reported at 92.5% average success on four real tasks with no robot data
ImportanceLocalEvidenceE2 unreplicated
With just 30 minutes of egocentric human video per task, you can train a manipulation policy that needs no robot data at all, achieving 92.5% average success on four real tasks — self-reported by Zhi Wang and six co-authors in HumanEgo. Previously, such policies relied on robot-collected data, which is costly and hard to transfer zero-shot to new robots, cameras, and environments.
The authors' self-reported measurements show: 30 minutes of human video per task, 92.5% average success across four real tasks, 41% higher than teleoperation of equal duration; the method converts egocentric human video into entity-level hand-object representations and trains a flow-matching policy; code and dataset are public.
Boundary: results await independent replication; the preprint was first posted May 24 and updated to v3 on September 14 (arXiv:2605.24934).
Sources: arXiv preprint ↗ | Code repository ↗ | Project page ↗