Reducing Intrinsic and Extrinsic Data Biases for Moment Localization with Natural Language
Jiong Yin, Liang Li, Jiehua Zhang, Chenggang Yan, Lei Zhang, Zunjie Zhu
摘要
Moment Localization with Natural Language (MLNL) aims to locate the target moment from an untrimmed video by a linguistic query. Recent works reveal the severe data bias problem in MLNL and point out that the multi-modal content may not be understood by fitting the timestamp distribution. In this paper, we study the data biases on the intrinsic and extrinsic aspects: the former is mainly caused by the ambiguity of the moment boundary and the information imbalance between input and output; The latter results from the long-tail distribution of moments in MLNL datasets. To alleviate this, we propose a hybrid multi-modal debiasing network with temporal consistency constraint for MLNL. Specifically, we first design the multi-temporal Transformer to mitigate the ambiguity of boundary by integrating frame-wise features into segment-wise and dynamically matching with moment boundaries. Then, we introduce the temporal consistency constraint that highlights the action information in complex moment content to overcome the intrinsic bias from information imbalance.Furthermore, we design the hybrid linguistic activating module with external knowledge to relieve the extrinsic bias, which introduces a prior guidance to focus the discriminative information from the tail samples. Extensive experiments on three public datasets demonstrate that our model outperforms the existing methods.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Debiased Teacher for Day-to-Night Domain Adaptive Object DetectionYiming Cui, Liang Li, Haibing Yin, Yuhan Gao 等ICCV 2025 · 被引用 2 次
- SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and SynopsesChaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi 等ACM MM 2024 · 被引用 2 次
- Reverse Distribution Based Video Moment Retrieval for Effective Bias EliminationLingdu Kong, Xiaochun Yang, Tieying Li, Bin Wang 等AAAI 2025
- Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie DubbingZhedong Zhang, Liang Li, Chenggang Yan, Chunshan Liu 等CVPR 2025
- Expert-Teacher-Student Collaborative Learning for Domain Adaptive Object DetectionYiming Cui, Liang Li, Haibing Yin, Yuhan Gao 等CVPR 2026
相关 Paper
- MS-DETR: Natural Language Video Localization with Sampling Moment-Moment InteractionJing Wang, Aixin Sun, Hao Zhang, Xiaoli LiACL 2023 · 被引用 13 次
- Structured Multi-Level Interaction Network for Video Moment Localization via Language QueryHao Wang, Zheng-Jun Zha, Liang Li, Dong Liu 等CVPR 2021
- Dual Path Interaction Network for Video Moment LocalizationHao Wang, Zheng-Jun Zha, Xuejin Chen, Zhiwei Xiong 等ACM MM 2020 · 被引用 69 次
- Multi-Stage Aggregated Transformer Network for Temporal Language Localization in VideosMingxing Zhang, Yang Yang, Xinghan Chen, Yanli Ji 等CVPR 2021
- Natural Language Video Localization with Learnable Moment ProposalsShaoning Xiao, Long Chen, Jian Shao, Yueting Zhuang 等EMNLP 2021 · 被引用 44 次
