Deconfounded Multimodal Learning for Spatio-temporal Video Grounding
Jiawei Wang, Zhanchang Ma, Da Cao, Yuquan Le, Junbin Xiao, Tat-Seng Chua
摘要
The task of spatio-temporal video grounding involves identifying the spatial and temporal regions in a video that correspond to the objects or actions described in a given textual description. However, current models used for spatio-temporal video grounding often rely heavily on spatio-temporal priors to make the predictions. As a result, they may suffer from spurious correlations and lack the ability to generalize well to new or diverse scenarios. To overcome this limitation, we introduce a deconfounded multimodal learning framework, which utilizes a structural causal model to treat dataset biases as a confounder and subsequently remove their confounding effect. Through this framework, we can perform causal intervention on the multimodal input and derive an unbiased estimation formula through the do-calculus technique. In order to tackle the challenge of diverse and often unobservable confounders, we further propose a novel retrieval-based approach with a causal mask mechanism. The proposed method leverages analogical reasoning to facilitate deconfounded learning and mitigate dataset biases, enabling unbiased spatio-temporal prediction without explicitly modeling the confounding factors. Extensive experiments on two challenging benchmarks have well verified the effectiveness and rationality of our proposed solution.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- AutoM3L: An Automated Multimodal Machine Learning Framework with Large Language ModelsDaqin Luo, Chengjian Feng, Yuxuan Nong, Yiqing ShenACM MM 2024 · 被引用 16 次
- Semantic Codebook Learning for Dynamic Recommendation ModelsZheqi Lv, Shaoxuan He, Tianyu Zhan, Shengyu Zhang 等ACM MM 2024 · 被引用 8 次
- Tackling Device Data Distribution Real-time Shift via Prototype-based Parameter EditingZheqi Lv, Wenqiao Zhang, Kairui Fu, Qi Tian 等ACM MM 2025
- Neural Causal Graph for Interpretable and Intervenable ClassificationJiawei Wang, Shaofei Lu, Da Cao, Dongyu Wang 等ICLR 2025
相关 Paper
- Interventional Video Grounding With Dual Contrastive LearningGuoshun Nan, Rui Qiao, Yao Xiao, Jun Liu 等CVPR 2021
- Boosting Temporal Sentence Grounding via Causal InferenceKefan Tang, Lihuo He, Jisheng Dang, Xinbo GaoACM MM 2025 · 被引用 2 次
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang 等SIGIR 2021 · 被引用 198 次
- CausalVTG: Towards Robust Video Temporal Grounding via Causal InferenceQiyi Wang, Senda Chen, Ying ShenNeurIPS 2025 · 被引用 1 次
- Contextual Debiasing for Visual Recognition with Causal MechanismsRuyang Liu, Hao Liu, Ge Li, Haodi Hou 等CVPR 2022 · 被引用 42 次
