Introducing Decomposed Causality with Spatiotemporal Object-Centric Representation for Video Classification
Yachong Zhang, Lei Meng, Shuo Xu, Zhuang Qi, Wei Wu, Lei Wu, Xiangxu Meng
摘要
Video classification requires event-level representations of objects and their interactions. Existing methods typically rely on data-driven approaches, which either learn such features from whole frames or object-centric visual regions. Therefore, the modeling of spatiotemporal interactions among objects is usually overlooked. To address this issue, this paper presents a Decomposition of Synergistic, Unique, and Redundant Causal Representations Learning (SurdCRL) model for video classification, which introduces a newly-proposed SURD causal theory to model the spatiotemporal features of both object dynamics and their in- and cross-frame interactions. Specifically, SurdCRL employs three modules to model the object-centric spatiotemporal dynamics using distinct types of causal components, where the first module Spatial-Temporal Entity Modeling decouples the frame into object and context entities, and employs a temporal message passing block to capture object state changes over time, generating spatiotemporal features as basic causal variables. Second, the Dual-Path Causal Inference module mitigates confounders among causal variables by front-door and back-door interventions, thus enabling the subsequent causal components to reflect their intrinsic effects. Finally, the Causal Composition and Selection module employs the compositional structure-aware attention to project the causal variables and their high-order interactions into the synergistic, unique, and redundant components. Experiments on two benchmarking datasets verify that SurdCRL better captures event-relevant object-centric representation by decomposing spatiotemporal object interactions into three types of causal components.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang 等ICLR 2023 · 被引用 753 次
- MViTv2: Improved Multiscale Vision Transformers for Classification and DetectionYanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam 等CVPR 2022 · 被引用 699 次
相关 Paper
- Modeling Event-level Causal Representation for Video ClassificationYuqing Wang, Lei Meng, Haokai Ma, Yuqing Wang 等ACM MM 2024 · 被引用 3 次
- Diversifying Spatial-Temporal Perception for Video Domain GeneralizationKun-Yu Lin, Jia-Run Du, Yipeng Gao, Jiaming Zhou 等NeurIPS 2023 · 被引用 27 次
- Interventional Video Relation DetectionYicong Li, Xun Yang, Xindi Shang, Tat-Seng ChuaACM MM 2021 · 被引用 61 次
- SEAL: Semantic Attention Learning for Long Video RepresentationLan Wang, Yujia Chen, Du Tran, Vishnu Naresh Boddeti 等CVPR 2025
- Interaction-Based Disentanglement of Entities for Object-Centric World ModelsAkihiro Nakano, Masahiro Suzuki, Yutaka MatsuoICLR 2023
