Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
Thong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy T. Nguyen, See-Kiong Ng, Anh Tuan Luu
Abstract
To equip artificial intelligence with a comprehensive understanding towards a temporal world, video and 4D panoptic scene graph generation abstracts visual data into nodes to represent entities and edges to capture temporal relations. Existing methods encode entity masks tracked across temporal dimensions (mask tubes), then predict their relations with temporal pooling operation, which does not fully utilize the motion indicative of the entities' relation. To overcome this limitation, we introduce a contrastive representation learning framework that focuses on motion pattern for temporal scene graph generation. Firstly, our framework encourages the model to learn close representations for mask tubes of similar subject-relation-object triplets. Secondly, we seek to push apart mask tubes from their temporally shuffled versions. Moreover, we also learn distant representations for mask tubes belonging to the same video but different triplets. Extensive experiments show that our motion-aware contrastive framework significantly improves state-of-the-art methods on both video and 4D datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb02b79a-17e6-465e-a11b-745385d75a2bCited by top-tier papers1
Ask how each one uses itBuilds on21
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Do Different Tracking Tasks Require Different Appearance Models?Zhongdao Wang, Hengshuang Zhao, Ya-Li Li, Shengjin Wang et al.NeurIPS 2021 · 107 citations
- Video K-Net: A Simple, Strong, and Unified Baseline for Video SegmentationXiangtai Li, Wenwei Zhang, Jiangmiao Pang, Kai Chen et al.CVPR 2022 · 71 citations
- Efficient Video Prediction via Sparsely Conditioned Flow MatchingAram Davtyan, Sepehr Sameni, Paolo FavaroICCV 2023 · 51 citations
- Dynamic Scene Graph Generation via Anticipatory Pre-trainingYiming Li, Xiaoshan Yang, Changsheng XuCVPR 2022 · 38 citations
Related papers
- Motion-Focused Contrastive Learning of Video Representations*Rui Li, Yiheng Zhang, Zhaofan Qiu, Ting Yao et al.ICCV 2021 · 37 citations
- Video Representation Learning with Graph Contrastive AugmentationJingran Zhang, Xing Xu, Fumin Shen, Yazhou Yao et al.ACM MM 2021 · 6 citations
- 4D Panoptic Scene Graph GenerationJingkang Yang, Jun Cen, Wenxuan Peng, Shuai Liu et al.NeurIPS 2023 · 33 citations
- Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation LearningMinghao Zhu, Xiao Lin, Ronghao Dang, Chengju Liu et al.ACM MM 2023 · 6 citations
- DyTed: Disentangled Representation Learning for Discrete-time Dynamic GraphKaike Zhang, Qi Cao, Gaolin Fang, Bingbing Xu et al.KDD 2023 · 28 citations
