Context-Aware Biaffine Localizing Network for Temporal Sentence Grounding
Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou, Yu Cheng, Wei Wei, Zichuan Xu, Yulai Xie
Abstract
This paper addresses the problem of temporal sentence grounding (TSG), which aims to identify the temporal boundary of a specific segment from an untrimmed video by a sentence query. Previous works either compare pre-defined candidate segments with the query and select the best one by ranking, or directly regress the boundary timestamps of the target segment. In this paper, we propose a novel localization framework that scores all pairs of start and end indices within the video simultaneously with a biaffine mechanism. In particular, we present a Contextaware Biaffine Localizing Network (CBLN) which incorporates both local and global contexts into features of each start/end position for biaffine-based localization. The local contexts from the adjacent frames help distinguish the visually similar appearance, and the global contexts from the entire video contribute to reasoning the temporal relation. Besides, we also develop a multi-modal self-attention module to provide fine-grained query-guided video representation for this biaffine strategy. Extensive experiments show that our CBLN significantly outperforms state-of-thearts on three public datasets (ActivityNet Captions, TACoS, and Charades-STA), demonstrating the effectiveness of the proposed localization framework. The code is available at https://github.com/liudaizong/CBLN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7edbfa48-7705-48af-8f6a-43140f38573cCited by top-tier papers56
- Negative Sample Matters: A Renaissance of Metric Learning for Temporal GroundingZhenzhi Wang, Limin Wang, Tao Wu, Tianhao Li et al.AAAI 2022 · 170 citations
- MomentDiff: Generative Video Moment Retrieval from Random to RealPandeng Li, Chen-Wei Xie, Hongtao Xie, Liming Zhao et al.NeurIPS 2023 · 113 citations
- Knowing Where to Focus: Event-aware Transformer for Video GroundingJinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon et al.ICCV 2023 · 103 citations
- MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio DescriptionsMattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron et al.CVPR 2022 · 84 citations
- Harmonizing Visual Text Comprehension and GenerationZhen Zhao, Jingqun Tang, Binghong Wu, Chunhui Lin et al.NeurIPS 2024 · 69 citations
Builds on8
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan et al.EMNLP 2020 · 387 citations
- Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware PredictionJingwen Wang, Lin Ma, Wenhao JiangAAAI 2020 · 206 citations
- Rethinking the Bottom-Up Framework for Query-Based Video LocalizationLong Chen, Chujie Lu, Siliang Tang, Jun Xiao et al.AAAI 2020 · 182 citations
- Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment LocalizationDaizong Liu, Xiaoye Qu, Xiao-Yang Liu, Jianfeng Dong et al.ACM MM 2020 · 115 citations
Related papers
- Fine-grained Iterative Attention Network for Temporal Language Localization in VideosXiaoye Qu, Pengwei Tang, Zhikang Zou, Yu Cheng et al.ACM MM 2020 · 92 citations
- Cascaded Prediction Network via Segment Tree for Temporal Video GroundingYang Zhao, Zhou Zhao, Zhu Zhang, Zhijie LinCVPR 2021
- Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal LearningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng et al.CVPR 2022 · 108 citations
- Temporal Sentence Grounding with Relevance Feedback in VideosJianfeng Dong, Xiaoman Peng, Daizong Liu, Xiaoye Qu et al.NeurIPS 2024 · 12 citations
- Relation-aware Video Reading Comprehension for Temporal Language GroundingJialin Gao, Xin Sun, Mengmeng Xu, Xi Zhou et al.EMNLP 2021 · 51 citations
