Mixup-Augmented Temporally Debiased Video Grounding with Content-Location Disentanglement
Xin Wang, Zihao Wu, Hong Chen, Xiaohan Lan, Wenwu Zhu
Abstract
Video Grounding (VG), has drawn widespread attention over the past few years, and numerous studies have been devoted to improving performance on various VG benchmarks. Nevertheless, the label annotation procedures in VG produce imbalanced query-moment-label distributions in the datasets, which severely deteriorate the learning model's capability of truly understanding the video contents. Existing works on debiased VG either focus on adjusting the learning model or conducting video-level augmentation, failing to handle the temporal bias issue caused by imbalanced query-moment-label distributions. In this paper, we propose a Disentangled Feature Mixup (DFM) framework for debiased VG, which is capable of performing unbiased grounding to tackle the temporal bias issue. Specifically, a feature-mixup augmentation strategy is designed to generate new (text, location) pairs with diverse temporal distributions via jointly augmenting the representation of text queries and the location labels. This strategy encourages making prediction based on more diverse data samples with balanced query-moment-label distributions. Furthermore, we also design a content-location disentanglement module to disentangle the representations of the temporal information and content information in videos, which is able to remove the spurious effect of temporal biases on video representation. Given that our proposed DFM framework conducts feature-level augmentation and disentanglement, it is model-agnostic and can be applied to most baselines simply yet effectively. Extensive experiments show that our proposed DFM framework is able to significantly outperform baseline models in various metrics under both independent identical distribution (i.i.d.) and out-of-distribution (o.o.d.) scenes, especially in scenarios with annotation distribution changes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b51ed01-03ae-4392-bb38-a819b85a4e89Cited by top-tier papers7
- LLM4DyG: Can Large Language Models Solve Spatial-Temporal Problems on Dynamic Graphs?Zeyang Zhang, Xin Wang, Ziwei Zhang, Haoyang Li et al.KDD 2024 · 32 citations
- Unsupervised Graph Neural Architecture Search with Disentangled Self-SupervisionZeyang Zhang, Xin Wang, Ziwei Zhang, Guangyao Shen et al.NeurIPS 2023 · 22 citations
- Disentangled Continual Graph Neural Architecture Search with Invariant Modular SupernetZeyang Zhang, Xin Wang, Yijian Qin, Hong Chen et al.ICML 2024 · 14 citations
- Boosting Temporal Sentence Grounding via Causal InferenceKefan Tang, Lihuo He, Jisheng Dang, Xinbo GaoACM MM 2025 · 2 citations
- Improving Compositional Generalization in Cross-Embodiment Learning via Mixture of Disentangled PrototypesRen Wang, Xin Wang, Tongtong Feng, Xinyue Gong et al.ACM MM 2025 · 2 citations
Builds on20
- Disentangled Graph Collaborative FilteringXiang Wang, Hongye Jin, An Zhang, Xiangnan He et al.SIGIR 2020 · 621 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 279 citations
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang et al.SIGIR 2021 · 198 citations
Related papers
- Curriculum Multi-Negative Augmentation for Debiased Video GroundingXiaohan Lan, Yitian Yuan, Hong Chen, Xin Wang et al.AAAI 2023 · 26 citations
- Embracing Uncertainty: Decoupling and De-Bias for Robust Temporal GroundingHao Zhou, Chongyang Zhang, Yan Luo, Yanjun Chen et al.CVPR 2021
- Counterfactually Augmented Event Matching for De-biased Temporal Sentence GroundingXun Jiang, Zhuoyuan Wei, Shenshen Li, Xing Xu et al.ACM MM 2024 · 10 citations
- Interventional Video Grounding With Dual Contrastive LearningGuoshun Nan, Rui Qiao, Yao Xiao, Jun Liu et al.CVPR 2021
- Reducing the Vision and Language Bias for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Wei HuACM MM 2022 · 52 citations
