Object Dynamics Distillation for Scene Decomposition and Representation
Qu Tang, Xiangyu Zhu, Zhen Lei, Zhaoxiang Zhang
Abstract
The ability to perceive scenes in terms of abstract entities is crucial for us toachieve higher-level intelligence. Recently, several methods have been proposedto learn object-centric representations of scenes with multiple objects, yet mostof which focus on static scenes. In this paper, we work on object dynamics andpropose Object Dynamics Distillation Network (ODDN), a framework that distillates explicit object dynamics (e.g., velocity) from sequential static representations. ODDN also builds a relation module to model object interactions. We verifyour approach on tasks of video reasoning and video prediction, which are two important evaluations for video understanding. The results show that the reasoningmodel with visual representations of ODDN performs better in answering reasoning questions around physical events in a video compared to the previous state-of-the-art methods. The distilled object dynamics also could be used to predictfuture video frames given two input frames, involving occlusion and objects collision. In addition, our architecture brings better segmentation quality and higherreconstruction accuracy.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 69ae0fdb-cd95-4334-abfd-441dd19825abCited by top-tier papers2
- Deconfounding Physical Dynamics with Global Causal Relation and Confounder Transmission for Counterfactual PredictionZongzhao Li, Xiangyu Zhu, Zhen Lei, Zhaoxiang ZhangAAAI 2022 · 6 citations
- Intrinsic Physical Concepts Discovery with Object-Centric Predictive ModelsQu Tang, Xiangyu Zhu, Zhen Lei, Zhaoxiang ZhangCVPR 2023
Related papers
- Relation Distillation Networks for Video Object DetectionJiajun Deng, Yingwei Pan, Ting Yao, Wengang Zhou et al.ICCV 2019 · 211 citations
- Ock: Unsupervised Dynamic Video Prediction With Object-Centric KinematicsYeon-Ji Song, Jaein Kim, Suhyung Choi, Jin-Hwa Kim et al.ICCV 2025 · 4 citations
- Multimodal Global Relation Knowledge Distillation for Egocentric Action AnticipationYi Huang, Xiaoshan Yang, Changsheng XuACM MM 2021 · 11 citations
- Generative Video Transformer: Can Objects be the Words?Yi-Fu Wu, Jaesik Yoon, Sungjin AhnICML 2021 · 37 citations
- Object-Centric Representation Learning with Generative Spatial-Temporal FactorizationNanbo Li, Muhammad Ahmed Raza, Wenbin Hu, Zhaole Sun et al.NeurIPS 2021 · 17 citations
