Object Dynamics Distillation for Scene Decomposition and Representation
Qu Tang, Xiangyu Zhu, Zhen Lei, Zhaoxiang Zhang
摘要
The ability to perceive scenes in terms of abstract entities is crucial for us toachieve higher-level intelligence. Recently, several methods have been proposedto learn object-centric representations of scenes with multiple objects, yet mostof which focus on static scenes. In this paper, we work on object dynamics andpropose Object Dynamics Distillation Network (ODDN), a framework that distillates explicit object dynamics (e.g., velocity) from sequential static representations. ODDN also builds a relation module to model object interactions. We verifyour approach on tasks of video reasoning and video prediction, which are two important evaluations for video understanding. The results show that the reasoningmodel with visual representations of ODDN performs better in answering reasoning questions around physical events in a video compared to the previous state-of-the-art methods. The distilled object dynamics also could be used to predictfuture video frames given two input frames, involving occlusion and objects collision. In addition, our architecture brings better segmentation quality and higherreconstruction accuracy.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Deconfounding Physical Dynamics with Global Causal Relation and Confounder Transmission for Counterfactual PredictionZongzhao Li, Xiangyu Zhu, Zhen Lei, Zhaoxiang ZhangAAAI 2022 · 被引用 6 次
- Intrinsic Physical Concepts Discovery with Object-Centric Predictive ModelsQu Tang, Xiangyu Zhu, Zhen Lei, Zhaoxiang ZhangCVPR 2023
相关 Paper
- Relation Distillation Networks for Video Object DetectionJiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou 等ICCV 2019 · 被引用 211 次
- Ock: Unsupervised Dynamic Video Prediction With Object-Centric KinematicsYeon-Ji Song, Jaein Kim, Suhyung Choi, Jin-Hwa Kim 等ICCV 2025 · 被引用 4 次
- Multimodal Global Relation Knowledge Distillation for Egocentric Action AnticipationYi Huang, Xiaoshan Yang, Changsheng XuACM MM 2021 · 被引用 11 次
- Generative Video Transformer: Can Objects be the Words?Yi-Fu Wu, Jaesik Yoon, Sungjin AhnICML 2021 · 被引用 37 次
- Object-Centric Representation Learning with Generative Spatial-Temporal FactorizationNanbo Li, Muhammad Ahmed Raza, Wenbin Hu, Zhaole Sun 等NeurIPS 2021 · 被引用 17 次
