Reducing Intrinsic and Extrinsic Data Biases for Moment Localization with Natural Language
Jiong Yin, Liang Li, Jiehua Zhang, Chenggang Yan, Lei Zhang, Zunjie Zhu
Abstract
Moment Localization with Natural Language (MLNL) aims to locate the target moment from an untrimmed video by a linguistic query. Recent works reveal the severe data bias problem in MLNL and point out that the multi-modal content may not be understood by fitting the timestamp distribution. In this paper, we study the data biases on the intrinsic and extrinsic aspects: the former is mainly caused by the ambiguity of the moment boundary and the information imbalance between input and output; The latter results from the long-tail distribution of moments in MLNL datasets. To alleviate this, we propose a hybrid multi-modal debiasing network with temporal consistency constraint for MLNL. Specifically, we first design the multi-temporal Transformer to mitigate the ambiguity of boundary by integrating frame-wise features into segment-wise and dynamically matching with moment boundaries. Then, we introduce the temporal consistency constraint that highlights the action information in complex moment content to overcome the intrinsic bias from information imbalance.Furthermore, we design the hybrid linguistic activating module with external knowledge to relieve the extrinsic bias, which introduces a prior guidance to focus the discriminative information from the tail samples. Extensive experiments on three public datasets demonstrate that our model outperforms the existing methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers5
- Debiased Teacher for Day-to-Night Domain Adaptive Object DetectionYiming Cui, Liang Li, Haibing Yin, Yuhan Gao et al.ICCV 2025 · 2 citations
- SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and SynopsesChaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi et al.ACM MM 2024 · 2 citations
- Reverse Distribution Based Video Moment Retrieval for Effective Bias EliminationLingdu Kong, Xiaochun Yang, Tieying Li, Bin Wang et al.AAAI 2025
- Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie DubbingZhedong Zhang, Liang Li, Chenggang Yan, Chunshan Liu et al.CVPR 2025
- Expert-Teacher-Student Collaborative Learning for Domain Adaptive Object DetectionYiming Cui, Liang Li, Haibing Yin, Yuhan Gao et al.CVPR 2026
Related papers
- MS-DETR: Natural Language Video Localization with Sampling Moment-Moment InteractionJing Wang, Aixin Sun, Hao Zhang, Xiaoli LiACL 2023 · 13 citations
- Structured Multi-Level Interaction Network for Video Moment Localization via Language QueryHao Wang, Zheng-Jun Zha, Liang Li, Dong Liu et al.CVPR 2021
- Dual Path Interaction Network for Video Moment LocalizationHao Wang, Zheng-Jun Zha, Xuejin Chen, Zhiwei Xiong et al.ACM MM 2020 · 69 citations
- Multi-Stage Aggregated Transformer Network for Temporal Language Localization in VideosMingxing Zhang, Yang Yang, Xinghan Chen, Yanli Ji et al.CVPR 2021
- Natural Language Video Localization with Learnable Moment ProposalsShaoning Xiao, Long Chen, Jian Shao, Yueting Zhuang et al.EMNLP 2021 · 44 citations
