Introducing Decomposed Causality with Spatiotemporal Object-Centric Representation for Video Classification
Yachong Zhang, Lei Meng, Shuo Xu, Zhuang Qi, Wei Wu, Lei Wu, Xiangxu Meng
Abstract
Video classification requires event-level representations of objects and their interactions. Existing methods typically rely on data-driven approaches, which either learn such features from whole frames or object-centric visual regions. Therefore, the modeling of spatiotemporal interactions among objects is usually overlooked. To address this issue, this paper presents a Decomposition of Synergistic, Unique, and Redundant Causal Representations Learning (SurdCRL) model for video classification, which introduces a newly-proposed SURD causal theory to model the spatiotemporal features of both object dynamics and their in- and cross-frame interactions. Specifically, SurdCRL employs three modules to model the object-centric spatiotemporal dynamics using distinct types of causal components, where the first module Spatial-Temporal Entity Modeling decouples the frame into object and context entities, and employs a temporal message passing block to capture object state changes over time, generating spatiotemporal features as basic causal variables. Second, the Dual-Path Causal Inference module mitigates confounders among causal variables by front-door and back-door interventions, thus enabling the subsequent causal components to reflect their intrinsic effects. Finally, the Causal Composition and Selection module employs the compositional structure-aware attention to project the causal variables and their high-order interactions into the synergistic, unique, and redundant components. Experiments on two benchmarking datasets verify that SurdCRL better captures event-relevant object-centric representation by decomposing spatiotemporal object interactions into three types of causal components.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4d28e37-74bd-4741-9ac1-560b6ea5043aBuilds on25
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
- MViTv2: Improved Multiscale Vision Transformers for Classification and DetectionYanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam et al.CVPR 2022 · 699 citations
Related papers
- Modeling Event-level Causal Representation for Video ClassificationYuqing Wang, Lei Meng, Haokai Ma, Yuqing Wang et al.ACM MM 2024 · 3 citations
- Diversifying Spatial-Temporal Perception for Video Domain GeneralizationKun-Yu Lin, Jia-Run Du, Yipeng Gao, Jiaming Zhou et al.NeurIPS 2023 · 27 citations
- Interventional Video Relation DetectionYicong Li, Xun Yang, Xindi Shang, Tat-Seng ChuaACM MM 2021 · 61 citations
- SEAL: Semantic Attention Learning for Long Video RepresentationLan Wang, Yujia Chen, Du Tran, Vishnu Naresh Boddeti et al.CVPR 2025
- Interaction-Based Disentanglement of Entities for Object-Centric World ModelsAkihiro Nakano, Masahiro Suzuki, Yutaka MatsuoICLR 2023
