RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment Generalization
LIU SONGMING, Bangguo Li, Kai Ma, Lingxuan Wu, Hengkai Tan, Xiao Ouyang, Hang Su, Jun Zhu
摘要
Vision-Language-Action (VLA) models hold promise for generalist robotics but currently struggle with data scarcity, architectural inefficiencies, and the inability to generalize across different hardware platforms. We introduce RDT2, a robotic foundation model built upon a 7B parameter VLM designed to enable zero-shot deployment on novel embodiments for open-vocabulary tasks. To achieve this, we collected one of the largest open-source robotic datasets-over 10, 000 hours of demonstrations in diverse families-using an enhanced, embodiment-agnostic Universal Manipulation Interface (UMI). Our approach employs a novel three-stage training recipe that aligns discrete linguistic knowledge with continuous control via Residual Vector Quantization (RVQ), flow-matching, and distillation for realtime inference. Consequently, RDT2 becomes one of the first models that simultaneously zeroshot generalizes to unseen objects, scenes, instructions, and even robotic platforms. Besides, it outperforms state-of-the-art baselines in dexterous, long-horizon, and dynamic downstream tasks like playing table tennis. See project page for more information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level CompositionJiahang Cao, Yize Huang, Hanzhong Guo, Qiang Zhang 等ICLR 2026 · 被引用 14 次
- PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied ManipulationLingxuan Wu, Zijian Zhu, Lizhong Wang, Chengyang Ying 等ICML 2026
它引用的顶会 Paper17
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo 等NeurIPS 2024 · 被引用 1,004 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang 等ICLR 2022 · 被引用 753 次
- Behavior Transformers: Cloning modes with one stoneNur Muhammad Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, Lerrel PintoNeurIPS 2022 · 被引用 470 次
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai 等ICML 2026 · 被引用 394 次
相关 Paper
- RDT-1B: a Diffusion Foundation Model for Bimanual ManipulationSongming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan 等ICLR 2025
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic ForgettingAsher J. Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky 等ICLR 2026 · 被引用 58 次
- VideoVLA: Video Generators Can Be Generalizable Robot ManipulatorsYichao Shen, Fangyun Wei, Zhiying Du, Yaobo Liang 等NeurIPS 2025 · 被引用 73 次
- H-RDT: Human Manipulation Enhanced Bimanual Robotic ManipulationHongzhe Bi, Lingxuan Wu, Tianwei Lin, Hengkai Tan 等AAAI 2026 · 被引用 25 次
- Interleave-VLA: Enhancing Robot Manipulation with Image-Text Interleaved InstructionsCunxin Fan, Xiaosong Jia, Yihang Sun, Yixiao Wang 等ICLR 2026 · 被引用 49 次
