H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
Hongzhe Bi, Lingxuan Wu, Tianwei Lin, Hengkai Tan, Zhizhong Su, Hang Su, Jun Zhu
摘要
Imitation learning for robotic manipulation faces a fundamental challenge: the scarcity of large-scale, high-quality robot demonstration data. Recent robotic foundation models often pre-train on cross-embodiment robot datasets to increase data scale, while they face significant limitations as the diverse morphologies and action spaces across different robot embodiments make unified training challenging. In this paper, we present H-RDT (Human to Robotics Diffusion Transformer), a novel approach that leverages human manipulation data to enhance robot manipulation capabilities. Our key insight is that large-scale egocentric human manipulation videos with paired 3D hand pose annotations provide rich behavioral priors that capture natural manipulation strategies and can benefit robotic policy learning. We introduce a two-stage training paradigm: (1) pre-training on largescale egocentric human manipulation data, and (2) crossembodiment fine-tuning on robot-specific data with modular action encoders and decoders. Built on a diffusion transformer architecture with 2B parameters, H-RDT uses flow matching to model complex action distributions. The modular design of action encoder and decoder components enables effective knowledge transfer from the unified human embodiment to diverse robot platforms through efficient finetuning. Extensive evaluations encompassing both simulation and real-world experiments, single-task and multitask scenarios, as well as few-shot learning and robustness assessments, demonstrate that H-RDT outperforms training from scratch and existing state-of-the-art methods, including π0 and RDT, achieving significant improvements of 13.9% and 40.5% over training from scratch in simulation and real-world experiments, respectively. The results validate our core hypothesis that human manipulation data can serve as a powerful foundation for learning bimanual robotic manipulation policies. See our project page for code and pretrained models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Motus: A Unified Latent Action World ModelHongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang 等CVPR 2026 · 被引用 271 次
- DreamDojo: A Real-Time Robot World Model from Large-Scale Human VideosShenyuan Gao, William Liang, Kaiyuan Zheng, Ayaan Malik 等ICML 2026 · 被引用 96 次
- Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent GuidanceYang Zhang, Chenwei Wang, Ouyang Lu, Yuan Zhao 等ICLR 2026 · 被引用 21 次
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action GenerationXiangkai Ma, Lekai Xing, Han Zhang, Wenzhong Li 等CVPR 2026 · 被引用 11 次
- CUBic: Coordinated Unified Bimanual Perception and Control FrameworkXingyu Wang, Pengxiang Ding, Jingkai Xu, Donglin Wang 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper13
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai 等ICML 2026 · 被引用 394 次
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric VideoRyan Hoque, Peide Huang, David J. Yoon, Mouli Sivapurapu 等ICLR 2026 · 被引用 248 次
相关 Paper
- RDT-1B: a Diffusion Foundation Model for Bimanual ManipulationSongming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan 等ICLR 2025
- Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained TransformersLirui Wang, Xinlei Chen, Jialiang Zhao, Kaiming HeNeurIPS 2024 · 被引用 208 次
- EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human DataRyan Punamiya, Dhruv Patel, Patcharapong Aphiwetsa, Pranav Kuppili 等NeurIPS 2025 · 被引用 40 次
- VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic ManipulationHanzhi Chen, Boyang Sun, Anran Zhang, Marc Pollefeys 等CVPR 2025
- Human2Robot: Learning Robot Actions from Paired Human-Robot VideosSicheng Xie, Haidong Cao, Zejia Weng, Zhen Xing 等AAAI 2026 · 被引用 15 次
