Learning Long-term Visual Dynamics with Region Proposal Interaction Networks
Haozhi Qi, Xiaolong Wang, Deepak Pathak, Yi Ma, Jitendra Malik
摘要
Learning long-term dynamics models is the key to understanding physical common sense. Most existing approaches on learning dynamics from visual input sidestep long-term predictions by resorting to rapid re-planning with short-term models. This not only requires such models to be super accurate but also limits them only to tasks where an agent can continuously obtain feedback and take action at each step until completion. In this paper, we aim to leverage the ideas from success stories in visual recognition tasks to build object representations that can capture inter-object and object-environment interactions over a long-range. To this end, we propose Region Proposal Interaction Networks (RPIN), which reason about each object's trajectory in a latent region-proposal feature space. Thanks to the simple yet effective object representation, our approach outperforms prior methods by a significant margin both in terms of prediction quality and their ability to plan for downstream tasks, and also generalize well to novel environments. Code, pre-trained models, and more visualization results are available at this https URL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- VDT: General-purpose Video Diffusion Transformers via Mask ModelingHaoyu Lu, Guoxing Yang, Nanyi Fei, Yuqi Huo 等ICLR 2024 · 被引用 117 次
- Neural Production SystemsAniket Didolkar, Anirudh Goyal, Nan Rosemary Ke, Charles Blundell 等NeurIPS 2021 · 被引用 94 次
- Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and LanguageMingyu Ding, Zhenfang Chen, Tao Du, Ping Luo 等NeurIPS 2021 · 被引用 90 次
- Joint Hand Motion and Interaction Hotspots Prediction from Egocentric VideosShaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong WangCVPR 2022 · 被引用 69 次
- Planning for Sample Efficient Imitation LearningZhao-Heng Yin, Weirui Ye, Qifeng Chen, Yang GaoNeurIPS 2022 · 被引用 32 次
它引用的顶会 Paper4
- Learning to Simulate Complex Physics with Graph NetworksAlvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying 等ICML 2020 · 被引用 1,439 次
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 被引用 322 次
- Compositional Video PredictionYufei Ye, Maneesh Singh, Abhinav Gupta, Shubham TulsianiICCV 2019 · 被引用 84 次
相关 Paper
- A Critical View of Vision-Based Long-Term Dynamics Prediction Under Environment MisalignmentHanchen Xie, Jiageng Zhu, Mahyar Khayatkhoei, Jiazhi Li 等ICML 2023 · 被引用 3 次
- Pushing It Out of the Way: Interactive Visual NavigationKuo-Hao Zeng, Luca Weihs, Ali Farhadi, Roozbeh MottaghiCVPR 2021
- Visual Grounding of Learned Physical ModelsYunzhu Li, Toru Lin, Kexin Yi, Daniel Bear 等ICML 2020 · 被引用 88 次
- HyperDynamics: Meta-Learning Object and Agent Dynamics with HypernetworksZhou Xian, Shamit Lal, Hsiao-Yu Tung, Emmanouil Antonios Platanios 等ICLR 2021 · 被引用 32 次
- Object Dynamics Distillation for Scene Decomposition and RepresentationQu Tang, Xiangyu Zhu, Zhen Lei, Zhaoxiang ZhangICLR 2022 · 被引用 7 次
