Modeling Event-level Causal Representation for Video Classification
Yuqing Wang, Lei Meng, Haokai Ma, Yuqing Wang, Haibei Huang, Xiangxu Meng
摘要
Classifying videos differs from that of images in the need to capture the information on what has happened, instead of what is in the frames. Conventional methods typically follow the data-driven approach, which uses transformer-based attention models to extract and aggregate the features of video frames as the representation of the entire video. However, this approach tends to extract the object information of frames and may face difficulties in classifying the classes talking about events, such as "fixing bicycle". To address this issue, This paper presents an Event-level Causal Representation Learning (ECRL) model for the spatio-temporal modeling of both the in-frame object interactions and their cross-frame temporal correlations. Specifically, ECRL first employs a Frame-to-Video Causal Modeling (F2VCM) module, which simultaneously builds the in-frame causal graph with the background and foreground information and models their cross-frame correlations to construct a video-level causal graph. Subsequently, a Causality-aware Event-level Representation Inference (CERI) module is introduced to eliminate the spurious correlations in contexts and objects via the back- and front-door interventions, respectively. The former involves visual context de-biasing to filter out background confounders, while the latter employs global-local causal attention to capture event-level visual information. Experimental results on two benchmarking datasets verified that ECRL may better capture the cross-frame correlations to describe videos in event-level features. The source codes have been released at https://github.com/wyqcrystal/ECRL.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Cross-Silo Feature Space Alignment for Federated Learning on Clients with Imbalanced DataZhuang Qi, Lei Meng, Zhaochuan Li, Han Hu 等AAAI 2025 · 被引用 39 次
- Curriculum Conditioned Diffusion for Multimodal RecommendationYimeng Yang, Haokai Ma, Lei Meng, Shuo Xu 等AAAI 2025 · 被引用 12 次
- Causal Inference over Visual-Semantic-Aligned Graph for Image ClassificationLei Meng, Xiangxian Li, Xiaoshuo Yan, Haokai Ma 等AAAI 2025 · 被引用 11 次
- Explicit Modeling of Causal Factors and Confounders for Image ClassificationWei Wu, Lei Meng, Zhuang Qi, Zixuan Li 等AAAI 2026
相关 Paper
- Introducing Decomposed Causality with Spatiotemporal Object-Centric Representation for Video ClassificationYachong Zhang, Lei Meng, Shuo Xu, Zhuang Qi 等AAAI 2026
- Class-level Structural Relation Modeling and Smoothing for Visual Representation LearningZitan Chen, Zhuang Qi, Xiao Cao, Xiangxian Li 等ACM MM 2023 · 被引用 10 次
- Cross-Modal Dual-Causal Learning for Long-Term Action RecognitionShaowu Xu, Xibin Jia, Junyu Gao, Qianmei Sun 等ACM MM 2025
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang 等SIGIR 2021 · 被引用 198 次
- Contrastive Learning of Image Representations with Cross-Video Cycle-ConsistencyHaiping Wu, Xiaolong WangICCV 2021 · 被引用 35 次
