Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding
Chaolei Tan, Zihang Lin, Jian-Fang Hu, Wei-Shi Zheng, Jianhuang Lai
Abstract
Video Paragraph Grounding (VPG) is an essential yet challenging task in vision-language understanding, which aims to jointly localize multiple events from an untrimmed video with a paragraph query description. One of the critical challenges in addressing this problem is to comprehend the complex semantic relations between visual and textual modalities. Previous methods focus on modeling the contextual information between the video and text from a single-level perspective (i.e., the sentence level), ignoring rich visual-textual correspondence relations at different semantic levels, e.g., the video-word and video-paragraph correspondence. To this end, we propose a novel Hierarchical Semantic Correspondence Network (HSCNet), which explores multi-level visual-textual correspondence by learning hierarchical semantic alignment and utilizes dense supervision by grounding diverse levels of queries. Specifically, we develop a hierarchical encoder that encodes the multi-modal inputs into semantics-aligned representations at different levels. To exploit the hierarchical semantic correspondence learned in the encoder for multi-level supervision, we further design a hierarchical decoder that progressively performs finer grounding for lower-level queries conditioned on higher-level semantics. Extensive experiments demonstrate the effectiveness of HSCNet and our method significantly outstrips the state-of-the-arts on two challenging benchmarks, i.e., ActivityNet-Captions and TACoS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cecd524d-1d82-4e62-8554-c54773795033Cited by top-tier papers8
- CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial VideosTrong-Thuan Nguyen, Pha A. Nguyen, Xin Li, Jackson David Cothren et al.NeurIPS 2024 · 13 citations
- TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement LearningTao Wu, Li Yang, Gen Zhan, Yabin ZHANG et al.CVPR 2026 · 7 citations
- HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video UnderstandingTrong-Thuan Nguyen, Pha A. Nguyen, Khoa LuuCVPR 2024 · 5 citations
- Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph GroundingChaolei Tan, Jianhuang Lai, Wei-Shi Zheng, Jian-Fang HuCVPR 2024 · 5 citations
- Learning Multi-Scale Video-Text Correspondence for Weakly Supervised Temporal Article GrondingWenjia Geng, Yong Liu, Lei Chen, Sujia Wang et al.AAAI 2024 · 3 citations
Builds on14
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan et al.EMNLP 2020 · 387 citations
- Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual ConceptsYan Zeng, Xinsong Zhang, Hang LiICML 2022 · 371 citations
- Boundary Proposal Network for Two-stage Natural Language Video LocalizationShaoning Xiao, Long Chen, Songyang Zhang, Wei Ji et al.AAAI 2021 · 186 citations
- Rethinking the Bottom-Up Framework for Query-Based Video LocalizationLong Chen, Chujie Lu, Siliang Tang, Jun Xiao et al.AAAI 2020 · 182 citations
Related papers
- Dense Events Grounding in VideoPeijun Bao, Qian Zheng, Yadong MuAAAI 2021 · 37 citations
- Semi-supervised Video Paragraph Grounding with Contrastive EncoderXun Jiang, Xing Xu, Jingran Zhang, Fumin Shen et al.CVPR 2022 · 42 citations
- HANet: Hierarchical Alignment Networks for Video-Text RetrievalPeng Wu, Xiangteng He, Mingqian Tang, Yiliang Lv et al.ACM MM 2021 · 62 citations
- Proposal-Free Video Grounding with Contextual Pyramid NetworkKun Li, Dan Guo, Meng WangAAAI 2021 · 138 citations
- Exploiting Auxiliary Caption for Video GroundingHongxiang Li, Meng Cao, Xuxin Cheng, Yaowei Li et al.AAAI 2024 · 16 citations
