Reducing the Vision and Language Bias for Temporal Sentence Grounding
Daizong Liu, Xiaoye Qu, Wei Hu
Abstract
Temporal sentence grounding (TSG) is an important yet challenging task in multimedia information retrieval. Although previous TSG methods have achieved decent performance, they tend to capture the selection biases of frequently appeared video-query pairs in the dataset rather than present robust multimodal reasoning abilities, especially for the rarely appeared pairs. In this paper, we study the above issue of selection biases and accordingly propose a Debiasing-TSG (D-TSG) model to filter and remove the negative biases in both vision and language modalities for enhancing the model generalization ability. Specifically, we propose to alleviate the issue from two perspectives: 1) Feature distillation. We built a multi-modal debiasing branch to firstly capture the vision and language biases, and then apply a bias identification module to explicitly recognize the true negative biases and remove them from the benign multi-modal representations. 2) Contrastive sample generation. We construct two types of negative samples to enforce the model to accurately learn the aligned multi-modal semantics and make complete semantic reasoning. We apply the proposed model to both commonly and rarely appeared TSG cases, and demonstrate its effectiveness by achieving the state-of-the-art performance on three benchmark datasets (ActivityNet Caption, TACoS, and Charades-STA).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5593a349-fa45-4e54-844f-20d4c268eb26Cited by top-tier papers22
- Knowing Where to Focus: Event-aware Transformer for Video GroundingJinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon et al.ICCV 2023 · 103 citations
- Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video RetrievalJianfeng Dong, Minsong Zhang, Zheng Zhang, Xianke Chen et al.ICCV 2023 · 35 citations
- Skimming, Locating, then Perusing: A Human-Like Framework for Natural Language Video LocalizationDaizong Liu, Wei HuACM MM 2022 · 33 citations
- Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using LanguageXiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou et al.AAAI 2024 · 30 citations
- Hypotheses Tree Building for One-Shot Temporal Sentence LocalizationDaizong Liu, Xiang Fang, Pan Zhou, Xing Di et al.AAAI 2023 · 29 citations
Builds on19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 279 citations
- Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware PredictionJingwen Wang, Lin Ma, Wenhao JiangAAAI 2020 · 206 citations
- Rethinking the Bottom-Up Framework for Query-Based Video LocalizationLong Chen, Chujie Lu, Siliang Tang, Jun Xiao et al.AAAI 2020 · 182 citations
Related papers
- Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in VideoZhaobo Qi, Yibo Yuan, Xiaowen Ruan, Shuhui Wang et al.AAAI 2024 · 17 citations
- Context-Aware Biaffine Localizing Network for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou et al.CVPR 2021
- Memory-Guided Semantic Learning Network for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Xing Di, Yu Cheng et al.AAAI 2022 · 83 citations
- Counterfactually Augmented Event Matching for De-biased Temporal Sentence GroundingXun Jiang, Zhuoyuan Wei, Shenshen Li, Xing Xu et al.ACM MM 2024 · 10 citations
- Boosting Temporal Sentence Grounding via Causal InferenceKefan Tang, Lihuo He, Jisheng Dang, Xinbo GaoACM MM 2025 · 2 citations
