Mixup-Augmented Temporally Debiased Video Grounding with Content-Location Disentanglement
Xin Wang, Zihao Wu, Hong Chen, Xiaohan Lan, Wenwu Zhu
摘要
Video Grounding (VG), has drawn widespread attention over the past few years, and numerous studies have been devoted to improving performance on various VG benchmarks. Nevertheless, the label annotation procedures in VG produce imbalanced query-moment-label distributions in the datasets, which severely deteriorate the learning model's capability of truly understanding the video contents. Existing works on debiased VG either focus on adjusting the learning model or conducting video-level augmentation, failing to handle the temporal bias issue caused by imbalanced query-moment-label distributions. In this paper, we propose a Disentangled Feature Mixup (DFM) framework for debiased VG, which is capable of performing unbiased grounding to tackle the temporal bias issue. Specifically, a feature-mixup augmentation strategy is designed to generate new (text, location) pairs with diverse temporal distributions via jointly augmenting the representation of text queries and the location labels. This strategy encourages making prediction based on more diverse data samples with balanced query-moment-label distributions. Furthermore, we also design a content-location disentanglement module to disentangle the representations of the temporal information and content information in videos, which is able to remove the spurious effect of temporal biases on video representation. Given that our proposed DFM framework conducts feature-level augmentation and disentanglement, it is model-agnostic and can be applied to most baselines simply yet effectively. Extensive experiments show that our proposed DFM framework is able to significantly outperform baseline models in various metrics under both independent identical distribution (i.i.d.) and out-of-distribution (o.o.d.) scenes, especially in scenarios with annotation distribution changes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- LLM4DyG: Can Large Language Models Solve Spatial-Temporal Problems on Dynamic Graphs?Zeyang Zhang, Xin Wang, Ziwei Zhang, Haoyang Li 等KDD 2024 · 被引用 32 次
- Unsupervised Graph Neural Architecture Search with Disentangled Self-SupervisionZeyang Zhang, Xin Wang, Ziwei Zhang, Guangyao Shen 等NeurIPS 2023 · 被引用 22 次
- Disentangled Continual Graph Neural Architecture Search with Invariant Modular SupernetZeyang Zhang, Xin Wang, Yijian Qin, Hong Chen 等ICML 2024 · 被引用 14 次
- Boosting Temporal Sentence Grounding via Causal InferenceKefan Tang, Lihuo He, Jisheng Dang, Xinbo GaoACM MM 2025 · 被引用 2 次
- Improving Compositional Generalization in Cross-Embodiment Learning via Mixture of Disentangled PrototypesRen Wang, Xin Wang, Tongtong Feng, Xinyue Gong 等ACM MM 2025 · 被引用 2 次
它引用的顶会 Paper20
- Disentangled Graph Collaborative FilteringXiang Wang, Hongye Jin, An Zhang, Xiangnan He 等SIGIR 2020 · 被引用 621 次
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 被引用 279 次
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang 等SIGIR 2021 · 被引用 198 次
相关 Paper
- Curriculum Multi-Negative Augmentation for Debiased Video GroundingXiaohan Lan, Yitian Yuan, Hong Chen, Xin Wang 等AAAI 2023 · 被引用 26 次
- Embracing Uncertainty: Decoupling and De-Bias for Robust Temporal GroundingHao Zhou, Chongyang Zhang, Yan Luo, Yanjun Chen 等CVPR 2021
- Counterfactually Augmented Event Matching for De-biased Temporal Sentence GroundingXun Jiang, Zhuoyuan Wei, Shenshen Li, Xing Xu 等ACM MM 2024 · 被引用 10 次
- Interventional Video Grounding With Dual Contrastive LearningGuoshun Nan, Rui Qiao, Yao Xiao, Jun Liu 等CVPR 2021
- Reducing the Vision and Language Bias for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Wei HuACM MM 2022 · 被引用 52 次
