MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation
Chengshu Li, Mengdi Xu, Arpit Bahety, Hang Yin, Yunfan Jiang, Huang Huang, Josiah Wong, Sujay Garlanka, Cem Gokmen, Ruohan Zhang, Weiyu Liu, Jiajun Wu
摘要
Imitation learning from large-scale, diverse human demonstrations has been shown to be effective for training robots, but collecting such data is costly and time-consuming. This challenge intensifies for multi-step bimanual mobile manipulation, where humans must teleoperate both the mobile base and two high-DoF arms. Prior X-Gen works have developed automated data generation frameworks for static (bimanual) manipulation tasks, augmenting a few human demos in simulation with novel scene configurations to synthesize large-scale datasets. However, prior works fall short for bimanual mobile manipulation tasks for two major reasons: 1) a mobile base introduces the problem of how to place the robot base to enable downstream manipulation (reachability) and 2) an active camera introduces the problem of how to position the camera to generate data for a visuomotor policy (visibility). To address these challenges, MoMaGen formulates data generation as a constrained optimization problem that satisfies hard constraints (e.g., reachability) while balancing soft constraints (e.g., visibility while navigation). This formulation generalizes across most existing automated data generation approaches and offers a principled foundation for developing future methods. We evaluate on four multi-step bimanual mobile manipulation tasks and find that MoMaGen enables the generation of much more diverse datasets than previous methods. As a result of the dataset diversity, we also show that the data generated by MoMaGen can be used to train successful imitation learning policies using a single source demo. Furthermore, the trained policy can be fine-tuned with a very small amount of real-world data (40 demos) to be succesfully deployed on real robotic hardware. More details are on our project page: momagen.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control InterfaceYujie Zhao, Hongwei Fan, Di Chen, Shengcong Chen 等CVPR 2026 · 被引用 8 次
- CUBic: Coordinated Unified Bimanual Perception and Control FrameworkXingyu Wang, Pengxiang Ding, Jingkai Xu, Donglin Wang 等CVPR 2026 · 被引用 1 次
- Scalable Trajectory Generation for Whole-Body Mobile ManipulationYida Niu, Xinhai Chang, Xin Liu, Ziyuan Jiao 等CVPR 2026
它引用的顶会 Paper4
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 被引用 911 次
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto 等NeurIPS 2020 · 被引用 833 次
- Counterfactual Data Augmentation using Locally Factored DynamicsSilviu Pitis, Elliot Creager, Animesh GargNeurIPS 2020 · 被引用 126 次
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation SkillsJiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling 等ICLR 2023 · 被引用 21 次
相关 Paper
- HumanoidGen: Data Generation for Bimanual Dexterous Manipulation via LLM ReasoningZhi Jing, Siyuan Yang, Jicong Ao, Ting Xiao 等NeurIPS 2025 · 被引用 23 次
- MoManipVLA: Transferring Vision-language-action Models for General Mobile ManipulationZhenyu Wu, Yuheng Zhou, Xiuwei Xu, Ziwei Wang 等CVPR 2025
- What Matters in Learning from Large-Scale Datasets for Robot ManipulationVaibhav Saxena, Matthew Bronars, Nadun Ranawaka Arachchige, Kuancheng Wang 等ICLR 2025
- AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Affordance CorrespondenceJiawei Zhang, Kaizhe Hu, Yingqian Huang, Yuanchen Ju 等CVPR 2026
- AdaManip: Adaptive Articulated Object Manipulation Environments and Policy LearningYuanfei Wang, Xiaojie Zhang, Ruihai Wu, Yu Li 等ICLR 2025
