Counterfactually Augmented Event Matching for De-biased Temporal Sentence Grounding
Xun Jiang, Zhuoyuan Wei, Shenshen Li, Xing Xu, Jingkuan Song, Heng Tao Shen
Abstract
Temporal Sentence Grounding (TSG), which aims to localize events in untrimmed videos with a given language query, has been widely studied in the last decades. However, recently researchers have demonstrated that previous approaches are severely limited in out-of-distribution generalization, thus proposing the De-biased TSG challenge which requires models to overcome weakness towards outlier test samples. In this paper, we design a novel framework, termed Counterfactually-Augmented Event Matching (CAEM), which incorporates counterfactual data augmentation to learn event-query joint representations to resist the training bias. Specifically, it consists of three components: (1) A Temporal Counterfactual Augmentation module that generates counterfactual video-text pairs by temporally delaying events in the untrimmed video, enhancing the model's capacity for counterfactual thinking. (2) An Event-Query Matching model that is used to learn joint representations and predict corresponding matching scores for each event candidate. (3) A Counterfact-Adaptive Framework (CAF) that incorporates the counterfactual consistency rules on the matching process of the same event-query pairs, furtherly mitigating the bias learned from training sets. We conduct thorough experiments on two widely used DTSG datasets, i.e., Charades-CD and ActivityNet-CD, to evaluate our proposed CAEM method. Extensive experimental results show our proposed CAEM method outperforms recent state-of-the-art methods on all datasets. Our implementation code is available at https://github.com/CFM-MSG/CAEM_Code.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4f9d9c9e-dd06-4c71-9d9b-dfa728d92e27Cited by top-tier papers3
- Implicit Counterfactual Learning for Audio-Visual SegmentationMingfeng Zha, Tianyu Li, Guoyin Wang, Peng Wang et al.ICCV 2025 · 3 citations
- Boosting Temporal Sentence Grounding via Causal InferenceKefan Tang, Lihuo He, Jisheng Dang, Xinbo GaoACM MM 2025 · 2 citations
- PHGC: Procedural Heterogeneous Graph Completion for Natural Language Task Verification in Egocentric VideosXun Jiang, Zhiyi Huang, Xing Xu, Jingkuan Song et al.CVPR 2025
Related papers
- Curriculum Multi-Negative Augmentation for Debiased Video GroundingXiaohan Lan, Yitian Yuan, Hong Chen, Xin Wang et al.AAAI 2023 · 26 citations
- Reducing the Vision and Language Bias for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Wei HuACM MM 2022 · 52 citations
- Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in VideoZhaobo Qi, Yibo Yuan, Xiaowen Ruan, Shuhui Wang et al.AAAI 2024 · 17 citations
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu et al.CVPR 2022 · 63 citations
- Phrase-Level Temporal Relationship Mining for Temporal Sentence LocalizationMinghang Zheng, Sizhe Li, Qingchao Chen, Yuxin Peng et al.AAAI 2023 · 26 citations
