Object-centric 3D Motion Field for Robot Learning from Human Videos
Zhao-Heng Yin, Sherry Yang, Pieter Abbeel
摘要
Learning robot control policies from human videos is a promising direction for scaling up robot learning. However, how to extract action knowledge (or action representations) from videos for policy learning remains a key challenge. Existing action representations such as video frames, pixelflow, and pointcloud flow have inherent limitations such as modeling complexity or loss of information. In this paper, we propose to use object-centric 3D motion field to represent actions for robot learning from human videos, and present a novel framework for extracting this representation from videos for zero-shot control. We introduce two novel components in its implementation. First, a novel training pipeline for training a "denoising" 3D motion field estimator to extract fine object 3D motions from human videos with noisy depth robustly. Second, a dense object-centric 3D motion field prediction architecture that favors both cross-embodiment transfer and policy generalization to background. We evaluate the system in real world setups. Experiments show that our method reduces 3D motion estimation error by over 50% compared to the latest method, achieve 55% average success rate in diverse tasks where prior approaches fail (≲ 10%), and can even acquire fine-grained manipulation skills like insertion.
Figure 1: We propose a novel framework for robot learning from human demonstration videos without relying on any robot-collected data. Our approach learns to control robots by extracting and modeling 3D object motion fields from RGBD human videos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- PointWorld: Scaling 3D World Models for In-The-Wild Robotic ManipulationWenlong Huang, Yu-Wei Chao, Arsalan Mousavian, Ming-Yu Liu 等CVPR 2026 · 被引用 87 次
- Translating Flow to Policy via Hindsight Online ImitationYitian Zheng, Zhangchen Ye, Weijun Dong, Shengjie Wang 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai 等NeurIPS 2023 · 被引用 742 次
相关 Paper
- VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic ManipulationHanzhi Chen, Boyang Sun, Anran Zhang, Marc Pollefeys 等CVPR 2025
- Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-TrainingHaoran He, Chenjia Bai, Ling Pan, Weinan Zhang 等NeurIPS 2024 · 被引用 38 次
- EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric FlowYixiang Chen, Peiyan Li, Yan Huang, Jiabing Yang 等ICCV 2025 · 被引用 2 次
- Learning to Act from Actionless Videos through Dense CorrespondencesPo-Chen Ko, Jiayuan Mao, Yilun Du, Shao-Hua Sun 等ICLR 2024 · 被引用 181 次
- GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement LearningKelin Yu, Sheng Zhang, Harshit Soora, Furong Huang 等ICCV 2025
