Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation
Yang Tian, Sizhe Yang, Jia Zeng, Ping Wang, Dahua Lin, Hao Dong, Jiangmiao Pang
2025年份
65顶会引用
摘要
Figure 1: In contrast to previous methods that (a) conduct end-to-end naive behavior cloning from large-scale robotic data or (b) use decoupled visual prediction and inverse dynamics models to set goals and guide actions, we present end-to-end Predictive Inverse Dynamics Models (PIDM) that closes the loop between vision and action. Seer, the model we built, surpasses previous states of the art and demonstrates consistent improvements over the ablated version without pre-training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper65
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelJinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu 等ICLR 2026 · 被引用 335 次
- Motus: A Unified Latent Action World ModelHongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang 等CVPR 2026 · 被引用 271 次
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang 等NeurIPS 2025 · 被引用 244 次
- Unified Vision-Language-Action ModelYuqi Wang, Xinghang Li, Wenxuan Wang, Junbo Zhang 等ICLR 2026 · 被引用 144 次
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action PoliciesZhixuan Liang, Yizhuo Li, Tianshuo Yang, CHENGYUE WU 等ICML 2026 · 被引用 86 次
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai 等NeurIPS 2023 · 被引用 742 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu 等ICLR 2024 · 被引用 375 次
相关 Paper
- When Does Predictive Inverse Dynamics Outperform Behavior Cloning?Lukas Schäfer, Pallavi Choudhury, Abdelhak Lemkhenter, Chris Lovett 等ICML 2026 · 被引用 3 次
- Disentangled Robot Learning via Separate Forward and Inverse Dynamics PretrainingWenyao Zhang, Bozhou Zhang, Zekun Qi, Wenjun Zeng 等ICLR 2026 · 被引用 18 次
- DynBridge: Bridging Imagination and Control through Interaction Dynamics for Robot ManipulationAlex Wang, Zhiwei Dong, Qicheng Bai, Chenshi Zhang 等CVPR 2026
- Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene EvolutionBozhou Zhang, Nan Song, Jingyu Li, Xiatian Zhu 等NeurIPS 2025 · 被引用 29 次
- Manipulate by Seeing: Creating Manipulation Controllers from Pre-Trained RepresentationsJianren Wang, Sudeep Dasari, Mohan Kumar Srirama, Shubham Tulsiani 等ICCV 2023 · 被引用 16 次
