Relation-aware Video Reading Comprehension for Temporal Language Grounding
Jialin Gao, Xin Sun, Mengmeng Xu, Xi Zhou, Bernard Ghanem
Abstract
Temporal language grounding in videos aims to localize the temporal span relevant to the given query sentence. Previous methods treat it either as a boundary regression task or a span extraction task. This paper will formulate temporal language grounding into video reading comprehension and propose a Relation-aware Network (RaNet) to address it. This framework aims to select a video moment choice from the predefined answer set with the aid of coarse-and-fine choice-query interaction and choice-choice relation construction. A choicequery interactor is proposed to match the visual and textual information simultaneously in sentence-moment and token-moment levels, leading to a coarse-and-fine cross-modal interaction. Moreover, a novel multi-choice relation constructor is introduced by leveraging graph convolution to capture the dependencies among video moment choices for the best choice selection. Extensive experiments on ActivityNet-Captions, TACoS, and Charades-STA demonstrate the effectiveness of our solution. Codes will be available at https: //github.com/Huntersxsx/RaNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc6cb891-8eb2-40d4-93b2-2c2bdf0654b2Cited by top-tier papers16
- MomentDiff: Generative Video Moment Retrieval from Random to RealPandeng Li, Chen-Wei Xie, Hongtao Xie, Liming Zhao et al.NeurIPS 2023 · 113 citations
- Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment RetrievalZhihang Liu, Jun Li, Hongtao Xie, Pandeng Li et al.AAAI 2024 · 49 citations
- End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang et al.ACL 2022 · 46 citations
- Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight DetectionYicheng Xiao, Zhuoyan Luo, Yong Liu, Yue Ma et al.CVPR 2024 · 43 citations
- You Need to Read Again: Multi-granularity Perception Network for Moment Retrieval in VideosXin Sun, Xuan Wang, Jialin Gao, Qiong Liu et al.SIGIR 2022 · 42 citations
Builds on9
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 279 citations
- Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware PredictionJingwen Wang, Lin Ma, Wenhao JiangAAAI 2020 · 206 citations
- Location-Aware Graph Convolutional Networks for Video Question AnsweringDeng Huang, Peihao Chen, Runhao Zeng, Qing Du et al.AAAI 2020 · 187 citations
Related papers
- Context-Aware Biaffine Localizing Network for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou et al.CVPR 2021
- Fine-grained Iterative Attention Network for Temporal Language Localization in VideosXiaoye Qu, Pengwei Tang, Zhikang Zou, Yu Cheng et al.ACM MM 2020 · 92 citations
- Local-Global Video-Text Interactions for Temporal GroundingJonghwan Mun, Minsu Cho, Bohyung HanCVPR 2020
- Temporal Sentence Grounding with Relevance Feedback in VideosJianfeng Dong, Xiaoman Peng, Daizong Liu, Xiaoye Qu et al.NeurIPS 2024 · 12 citations
- Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment LocalizationDaizong Liu, Xiaoye Qu, Xiao-Yang Liu, Jianfeng Dong et al.ACM MM 2020 · 115 citations
