Proposal-Free Video Grounding with Contextual Pyramid Network
Kun Li, Dan Guo, Meng Wang
摘要
The challenge of video grounding -localizing activities in an untrimmed video via a natural language query -is to tackle the semantics of vision and language consistently along the temporal dimension. Most existing proposal-based methods are trapped by computational cost with extensive candidate proposals. In this paper, we propose a novel proposalfree framework named Contextual Pyramid Network (CP-Net) to investigate multi-scale temporal correlation in the video. Specifically, we propose a pyramid network to extract 2D contextual correlation maps at different temporal scales (T * T , T 2 * T 2 , T 4 * T 4 ), where the 2D correlation map (past → current & current ← future) is designed to model all the relations of any two moments in the video. In other words, CPNet progressively replenishes the temporal contexts and refines the location of queried activity by enlarging the temporal receptive fields. Finally, we implement a temporal self-attentive regression (i.e., proposal-free regression) to predict the activity boundary from the above hierarchical context-aware 2D correlation maps. Extensive experiments on ActivityNet Captions, Charades-STA, and TACoS datasets demonstrate that our approach outperforms state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- EulerMormer: Robust Eulerian Motion Magnification via Dynamic Filtering within TransformerFei Wang, Dan Guo, Kun Li, Meng WangAAAI 2024 · 被引用 49 次
- Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment RetrievalZhihang Liu, Jun Li, Hongtao Xie, Pandeng Li 等AAAI 2024 · 被引用 49 次
- Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight DetectionYicheng Xiao, Zhuoyan Luo, Yong Liu, Yue Ma 等CVPR 2024 · 被引用 43 次
- You Need to Read Again: Multi-granularity Perception Network for Moment Retrieval in VideosXin Sun, Xuan Wang, Jialin Gao, Qiong Liu 等SIGIR 2022 · 被引用 42 次
- Semi-supervised Video Paragraph Grounding with Contrastive EncoderXun Jiang, Xing Xu, Jingran Zhang, Fumin Shen 等CVPR 2022 · 被引用 42 次
它引用的顶会 Paper9
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 被引用 279 次
- Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware PredictionJingwen Wang, Lin Ma, Wenhao JiangAAAI 2020 · 被引用 206 次
- Rethinking the Bottom-Up Framework for Query-Based Video LocalizationLong Chen, Chujie Lu, Siliang Tang, Jun Xiao 等AAAI 2020 · 被引用 182 次
- Tree-Structured Policy Based Progressive Reinforcement Learning for Temporally Language Grounding in VideoJie Wu, Guanbin Li, Si Liu, Liang LinAAAI 2020 · 被引用 117 次
相关 Paper
- Cascaded Prediction Network via Segment Tree for Temporal Video GroundingYang Zhao, Zhou Zhao, Zhu Zhang, Zhijie LinCVPR 2021
- Boundary Proposal Network for Two-stage Natural Language Video LocalizationShaoning Xiao, Long Chen, Songyang Zhang, Wei Ji 等AAAI 2021 · 被引用 186 次
- Context-Aware Biaffine Localizing Network for Temporal Sentence GroundingDaizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 等CVPR 2021
- Hierarchical Semantic Correspondence Networks for Video Paragraph GroundingChaolei Tan, Zihang Lin, Jian-Fang Hu, Wei-Shi Zheng 等CVPR 2023
- Fine-grained Iterative Attention Network for Temporal Language Localization in VideosXiaoye Qu, Pengwei Tang, Zhikang Zou, Yu Cheng 等ACM MM 2020 · 被引用 92 次
