Deconfounded Video Moment Retrieval with Causal Intervention
Xun Yang, Fuli Feng, Wei Ji, Meng Wang, Tat-Seng Chua
摘要
We tackle the task of video moment retrieval (VMR), which aims to localize a specific moment in a video according to a textual query. Existing methods primarily model the matching relationship between query and moment by complex cross-modal interactions. Despite their effectiveness, current models mostly exploit dataset biases while ignoring the video content, thus leading to poor generalizability. We argue that the issue is caused by the hidden confounder in VMR, i.e., temporal location of moments, that spuriously correlates the model input and prediction. How to design robust matching models against the temporal location biases is crucial but, as far as we know, has not been studied yet for VMR.
To fill the research gap, we propose a causality-inspired VMR framework that builds structural causal model to capture the true effect of query and video content on the prediction. Specifically, we develop a Deconfounded Cross-modal Matching (DCM) method to remove the confounding effects of moment location. It first disentangles moment representation to infer the core feature of visual content, and then applies causal intervention on the disentangled multimodal input based on backdoor adjustment, which forces the model to fairly incorporate each possible location of the target into consideration. Extensive experiments clearly show that our approach can achieve significant improvement over the state-of-theart methods in terms of both accuracy and generalization (Codes: https://github.com/Xun-Yang/Causal_Video_Moment_Retrieval).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper65
- MomentDiff: Generative Video Moment Retrieval from Random to RealPandeng Li, Chen-Wei Xie, Hongtao Xie, Liming Zhao 等NeurIPS 2023 · 被引用 113 次
- Knowing Where to Focus: Event-aware Transformer for Video GroundingJinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon 等ICCV 2023 · 被引用 103 次
- Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose EstimationZhenguang Liu, Runyang Feng, Haoming Chen, Shuang Wu 等CVPR 2022 · 被引用 76 次
- Should Graph Convolution Trust Neighbors? A Simple Causal Inference MethodFuli Feng, Weiran Huang, Xiangnan He, Xin Xin 等SIGIR 2021 · 被引用 66 次
- Partially Relevant Video RetrievalJianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang 等ACM MM 2022 · 被引用 65 次
它引用的顶会 Paper17
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
- Causal Intervention for Weakly-Supervised Semantic SegmentationDong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua 等NeurIPS 2020 · 被引用 563 次
- Interventional Few-Shot LearningZhongqi Yue, Hanwang Zhang, Qianru Sun, Xian-Sheng HuaNeurIPS 2020 · 被引用 284 次
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 被引用 279 次
- Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware PredictionJingwen Wang, Lin Ma, Wenhao JiangAAAI 2020 · 被引用 206 次
相关 Paper
- Deconfounded Multimodal Learning for Spatio-temporal Video GroundingJiawei Wang, Zhanchang Ma, Da Cao, Yuquan Le 等ACM MM 2023 · 被引用 7 次
- Interventional Video Grounding With Dual Contrastive LearningGuoshun Nan, Rui Qiao, Yao Xiao, Jun Liu 等CVPR 2021
- Semantics-Enriched Cross-Modal Alignment for Complex-Query Video Moment RetrievalXingyu Shen, Xiang Zhang, Xun Yang, Yibing Zhan 等ACM MM 2023 · 被引用 9 次
- Counterfactual Cross-modality Reasoning for Weakly Supervised Video Moment LocalizationZezhong Lv, Bing Su, Ji-Rong WenACM MM 2023 · 被引用 23 次
- Video Corpus Moment Retrieval with Contrastive LearningHao Zhang, Aixin Sun, Wei Jing, Guoshun Nan 等SIGIR 2021 · 被引用 88 次
