Phrase-Level Temporal Relationship Mining for Temporal Sentence Localization
Minghang Zheng, Sizhe Li, Qingchao Chen, Yuxin Peng, Yang Liu
摘要
In this paper, we address the problem of video temporal sentence localization, which aims to localize a target moment from videos according to a given language query. We observe that existing models suffer from a sheer performance drop when dealing with simple phrases contained in the sentence. It reveals the limitation that existing models only capture the annotation bias of the datasets but lack sufficient understanding of the semantic phrases in the query. To address this problem, we propose a phrase-level Temporal Relationship Mining (TRM) framework employing the temporal relationship relevant to the phrase and the whole sentence to have a better understanding of each semantic entity in the sentence. Specifically, we use phrase-level predictions to refine the sentence-level prediction, and use Multiple Instance Learning to improve the quality of phrase-level predictions. We also exploit the consistency and exclusiveness constraints of phrase-level and sentence-level predictions to regularize the training process, thus alleviating the ambiguity of each phrase prediction. The proposed approach sheds light on how machines can understand detailed phrases in a sentence and their compositions in their generality rather than learning the annotation biases. Experiments on the ActivityNet Captions and Charades-STA datasets show the effectiveness of our method on both phrase and sentence temporal localization and enable better model interpretability and generalization when dealing with unseen compositions of seen concepts. Code can be found at https://github.com/minghangz/TRM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Generating Structured Pseudo Labels for Noise-resistant Zero-shot Video Sentence LocalizationMinghang Zheng, Shaogang Gong, Hailin Jin, Yuxin Peng 等ACL 2023 · 被引用 18 次
- Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in VideoZhaobo Qi, Yibo Yuan, Xiaowen Ruan, Shuhui Wang 等AAAI 2024 · 被引用 17 次
- ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual GroundingMinghang Zheng, Jiahua Zhang, Qingchao Chen, Yuxin Peng 等ACM MM 2024 · 被引用 5 次
- Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text UnderstandingJongbhin Woo, Hyeonggon Ryu, Youngjoon Jang, Jae-Won Cho 等ACM MM 2024 · 被引用 3 次
- Hierarchical Event Memory for Accurate and Low-Latency Online Video Temporal GroundingMinghang Zheng, Yuxin Peng, Benyuan Sun, Yi Yang 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper16
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
- Boundary Proposal Network for Two-stage Natural Language Video LocalizationShaoning Xiao, Long Chen, Songyang Zhang, Wei Ji 等AAAI 2021 · 被引用 186 次
- Semantic Grouping Network for Video CaptioningHobin Ryu, Sunghun Kang, Haeyong Kang, Chang D. YooAAAI 2021 · 被引用 160 次
- Fast Video Moment RetrievalJunyu Gao, Changsheng XuICCV 2021 · 被引用 132 次
- Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal LearningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng 等CVPR 2022 · 被引用 108 次
相关 Paper
- Cross-Sentence Temporal and Semantic Relations in Video Activity LocalisationJiabo Huang, Yang Liu, Shaogang Gong, Hailin JinICCV 2021 · 被引用 77 次
- Weakly-Supervised Video Moment Retrieval via Semantic Completion NetworkZhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang 等AAAI 2020 · 被引用 170 次
- Context-Aware Biaffine Localizing Network for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 等CVPR 2021
- Local-Global Video-Text Interactions for Temporal GroundingJonghwan Mun, Minsu Cho, Bohyung HanCVPR 2020
- Unsupervised Temporal Video Grounding with Deep Semantic ClusteringDaizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di 等AAAI 2022 · 被引用 52 次
