Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment Localization
Daizong Liu, Xiaoye Qu, Xiao-Yang Liu, Jianfeng Dong, Pan Zhou, Zichuan Xu
Abstract
Query-based moment localization is a new task that localizes the best matched segment in an untrimmed video according to a given sentence query. In this localization task, one should pay more attention to thoroughly mine visual and linguistic information. To this end, we propose a novel Cross-and Self-Modal Graph Attention Network (CSMGAN) that recasts this task as a process of iterative messages passing over a joint graph. Specifically, the joint graph consists of Cross-Modal relation Graph (CMG) and Self-Modal relation Graph (SMG), where frames and words are represented as nodes, and the relations between cross-and self-modal node pairs are described by an attention mechanism. Through parametric message passing, CMG highlights relevant instances across video and sentence, and then SMG models the pairwise relation inside each modality for frame (word) correlating. With multiple layers of such a joint graph, our CSMGAN is able to effectively capture high-order interactions between two modalities, thus enabling a further precise localization. Besides, to better comprehend the contextual details in the query, we develop a hierarchical sentence encoder to enhance the query understanding. Extensive experiments on two public datasets demonstrate the effectiveness of our proposed model, and GCSMAN significantly outperforms the state-of-the-arts. The code is available at https://github.com/liudaizong/CSMGAN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a40cece-5371-48ae-a4f7-4944e4c22015Cited by top-tier papers49
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang et al.SIGIR 2021 · 198 citations
- Fast Video Moment RetrievalJunyu Gao, Changsheng XuICCV 2021 · 132 citations
- Knowing Where to Focus: Event-aware Transformer for Video GroundingJinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon et al.ICCV 2023 · 103 citations
- MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio DescriptionsMattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron et al.CVPR 2022 · 84 citations
- Memory-Guided Semantic Learning Network for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Xing Di, Yu Cheng et al.AAAI 2022 · 83 citations
Builds on3
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware PredictionJingwen Wang, Lin Ma, Wenhao JiangAAAI 2020 · 206 citations
- Rethinking the Bottom-Up Framework for Query-Based Video LocalizationLong Chen, Chujie Lu, Siliang Tang, Jun Xiao et al.AAAI 2020 · 182 citations
Related papers
- Fine-grained Iterative Attention Network for Temporal Language Localization in VideosXiaoye Qu, Pengwei Tang, Zhikang Zou, Yu Cheng et al.ACM MM 2020 · 92 citations
- Context-Aware Biaffine Localizing Network for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou et al.CVPR 2021
- Modal-specific Pseudo Query Generation for Video Corpus Moment RetrievalMinjoon Jung, Seongho Choi, Joochan Kim, Jin-Hwa Kim et al.EMNLP 2022 · 11 citations
- STRONG: Spatio-Temporal Reinforcement Learning for Cross-Modal Video Moment LocalizationDa Cao, Yawen Zeng, Meng Liu, Xiangnan He et al.ACM MM 2020 · 47 citations
- Structured Multi-Level Interaction Network for Video Moment Localization via Language QueryHao Wang, Zheng-Jun Zha, Liang Li, Dong Liu et al.CVPR 2021
