Curriculum Multi-Negative Augmentation for Debiased Video Grounding
Xiaohan Lan, Yitian Yuan, Hong Chen, Xin Wang, Zequn Jie, Lin Ma, Zhi Wang, Wenwu Zhu
Abstract
Video Grounding (VG) aims to locate the desired segment from a video given a sentence query. Recent studies have found that current VG models are prone to over-rely the groundtruth moment annotation distribution biases in the training set. To discourage the standard VG model's behavior of exploiting such temporal annotation biases and improve the model generalization ability, we propose multiple negative augmentations in a hierarchical way, including cross-video augmentations from clip-/video-level, and self-shuffled augmentations with masks. These augmentations can effectively diversify the data distribution so that the model can make more reasonable predictions instead of merely fitting the temporal biases. However, directly adopting such data augmentation strategy may inevitably carry some noise shown in our cases, since not all of the handcrafted augmentations are semantically irrelevant to the groundtruth video. To further denoise and improve the grounding accuracy, we design a multi-stage curriculum strategy to adaptively train the standard VG model from easy to hard negative augmentations. Experiments on newly collected Charades-CD and ActivityNet-CD datasets demonstrate our proposed strategy can improve the performance of the base model on both i.i.d and o.o.d scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e88e62f7-5f82-455d-aab0-d783ec3bb825Cited by top-tier papers12
- Curriculum Co-disentangled Representation Learning across Multiple Environments for Social RecommendationXin Wang, Zirui Pan, Yuwei Zhou, Hong Chen et al.ICML 2023 · 31 citations
- Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in VideoZhaobo Qi, Yibo Yuan, Xiaowen Ruan, Shuhui Wang et al.AAAI 2024 · 17 citations
- CurBench: Curriculum Learning BenchmarkYuwei Zhou, Zirui Pan, Xin Wang, Hong Chen et al.ICML 2024 · 11 citations
- Mixup-Augmented Temporally Debiased Video Grounding with Content-Location DisentanglementXin Wang, Zihao Wu, Hong Chen, Xiaohan Lan et al.ACM MM 2023 · 9 citations
- Neighbor Does Matter: Curriculum Global Positive-Negative Sampling for Vision-Language Pre-trainingBin Huang, Feng He, Qi Wang, Hong Chen et al.ACM MM 2024 · 2 citations
Builds on13
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang et al.SIGIR 2021 · 198 citations
- Tree-Structured Policy Based Progressive Reinforcement Learning for Temporally Language Grounding in VideoJie Wu, Guanbin Li, Si Liu, Liang LinAAAI 2020 · 117 citations
- Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal LearningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng et al.CVPR 2022 · 108 citations
- MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio DescriptionsMattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron et al.CVPR 2022 · 84 citations
Related papers
- Counterfactually Augmented Event Matching for De-biased Temporal Sentence GroundingXun Jiang, Zhuoyuan Wei, Shenshen Li, Xing Xu et al.ACM MM 2024 · 10 citations
- Weakly Supervised Temporal Sentence Grounding with Uncertainty-Guided Self-trainingYifei Huang, Lijin Yang, Yoichi SatoCVPR 2023
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu et al.CVPR 2022 · 63 citations
- Embracing Uncertainty: Decoupling and De-Bias for Robust Temporal GroundingHao Zhou, Chongyang Zhang, Yan Luo, Yanjun Chen et al.CVPR 2021
- Reducing the Vision and Language Bias for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Wei HuACM MM 2022 · 52 citations
