MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding
Tongtong Cheng, Rongzhen Li, Yixin Xiong, Tao Zhang, Jing Wang, Kai Liu
Abstract
Accurate driving behavior recognition and reasoning are critical for autonomous driving video understanding. However, existing methods often tend to dig out the shallow causal, fail to address spurious correlations across modalities, and ignore the ego-vehicle level causality modeling. To overcome these limitations, we propose a novel Multimodal Causal Analysis Model (MCAM) that constructs latent causal structures between visual and language modalities. Firstly, we design a multi-level feature extractor to capture long-range dependencies. Secondly, we design a causal analysis module that dynamically models driving scenarios using a directed acyclic graph (DAG) of driving states. Thirdly, we utilize a vision-language transformer to align critical visual features with their corresponding linguistic expressions. Extensive experiments on the BDD-X, and CoVLA datasets demonstrate that MCAM achieves SOTA performance in visual-language causal relationship learning. Furthermore, the model exhibits superior capability in capturing causal characteristics within video sequences, showcasing its effectiveness for autonomous driving applications. The code is available at https://github.com/SixCorePeach/MCAM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d85e7af-8394-48ac-bdbf-c309727b713fBuilds on15
- SwinBERT: End-to-End Transformers with Sparse Attention for Video CaptioningKevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed et al.CVPR 2022 · 263 citations
- Spatial-Temporal Transformer for Dynamic Scene Graph GenerationYuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn et al.ICCV 2021 · 163 citations
- Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight DetectionYicheng Xiao, Zhuoyan Luo, Yong Liu, Yue Ma et al.CVPR 2024 · 43 citations
- Tem-adapter: Adapting Image-Text Pretraining for Video Question AnswerGuangyi Chen, Xiao Liu, Guangrun Wang, Kun Zhang et al.ICCV 2023 · 32 citations
- Visual Causal Scene Refinement for Video Question AnsweringYushen Wei, Yang Liu, Hong Yan, Guanbin Li et al.ACM MM 2023 · 31 citations
Related papers
- Cross-Modal Dual-Causal Learning for Long-Term Action RecognitionShaowu Xu, Xibin Jia, Junyu Gao, Qianmei Sun et al.ACM MM 2025
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action ModelXingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma et al.AAAI 2026 · 119 citations
- ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous DrivingQihang Peng, Xuesong Chen, Chenye Yang, Shaoshuai Shi et al.CVPR 2026 · 10 citations
- Unified Vision-Language-Action ModelYuqi Wang, Xinghang Li, Wenxuan Wang, Junbo Zhang et al.ICLR 2026 · 144 citations
- LLCP: Learning Latent Causal Processes for Reasoning-based Video Question AnswerGuangyi Chen, Yuke Li, Xiao Liu, Zijian Li et al.ICLR 2024 · 5 citations
