Modeling Event-level Causal Representation for Video Classification
Yuqing Wang, Lei Meng, Haokai Ma, Yuqing Wang, Haibei Huang, Xiangxu Meng
Abstract
Classifying videos differs from that of images in the need to capture the information on what has happened, instead of what is in the frames. Conventional methods typically follow the data-driven approach, which uses transformer-based attention models to extract and aggregate the features of video frames as the representation of the entire video. However, this approach tends to extract the object information of frames and may face difficulties in classifying the classes talking about events, such as "fixing bicycle". To address this issue, This paper presents an Event-level Causal Representation Learning (ECRL) model for the spatio-temporal modeling of both the in-frame object interactions and their cross-frame temporal correlations. Specifically, ECRL first employs a Frame-to-Video Causal Modeling (F2VCM) module, which simultaneously builds the in-frame causal graph with the background and foreground information and models their cross-frame correlations to construct a video-level causal graph. Subsequently, a Causality-aware Event-level Representation Inference (CERI) module is introduced to eliminate the spurious correlations in contexts and objects via the back- and front-door interventions, respectively. The former involves visual context de-biasing to filter out background confounders, while the latter employs global-local causal attention to capture event-level visual information. Experimental results on two benchmarking datasets verified that ECRL may better capture the cross-frame correlations to describe videos in event-level features. The source codes have been released at https://github.com/wyqcrystal/ECRL.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 48cc320f-6b0c-4d16-a227-2f599d766d11Cited by top-tier papers4
- Cross-Silo Feature Space Alignment for Federated Learning on Clients with Imbalanced DataZhuang Qi, Lei Meng, Zhaochuan Li, Han Hu et al.AAAI 2025 · 39 citations
- Curriculum Conditioned Diffusion for Multimodal RecommendationYimeng Yang, Haokai Ma, Lei Meng, Shuo Xu et al.AAAI 2025 · 12 citations
- Causal Inference over Visual-Semantic-Aligned Graph for Image ClassificationLei Meng, Xiangxian Li, Xiaoshuo Yan, Haokai Ma et al.AAAI 2025 · 11 citations
- Explicit Modeling of Causal Factors and Confounders for Image ClassificationWei Wu, Lei Meng, Zhuang Qi, Zixuan Li et al.AAAI 2026
Related papers
- Introducing Decomposed Causality with Spatiotemporal Object-Centric Representation for Video ClassificationYachong Zhang, Lei Meng, Shuo Xu, Zhuang Qi et al.AAAI 2026
- Class-level Structural Relation Modeling and Smoothing for Visual Representation LearningZitan Chen, Zhuang Qi, Xiao Cao, Xiangxian Li et al.ACM MM 2023 · 10 citations
- Cross-Modal Dual-Causal Learning for Long-Term Action RecognitionShaowu Xu, Xibin Jia, Junyu Gao, Qianmei Sun et al.ACM MM 2025
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang et al.SIGIR 2021 · 198 citations
- Contrastive Learning of Image Representations with Cross-Video Cycle-ConsistencyHaiping Wu, Xiaolong WangICCV 2021 · 35 citations
