Maskable Retentive Network for Video Moment Retrieval
Jingjing Hu, Dan Guo, Kun Li, Zhan Si, Xun Yang, Meng Wang
摘要
Video Moment Retrieval (MR) tasks involve predicting the moment described by a given natural language or spoken language query in an untrimmed video. In this paper, we propose a novel Maskable Retentive Network (MRNet) to address two key challenges in MR tasks: cross-modal guidance and video sequence modeling. Our approach introduces a new retention mechanism into the multimodal Transformer architecture, incorporating modality-specific attention modes. Specifically, we employ the Unlimited Attention for language-related attention regions to maximize cross-modal mutual guidance. Then, we introduce the Maskable Retention for video-only attention region to enhance video sequence modeling, that is, recognizing two crucial characteristics of video sequences: 1) bidirectional, decaying, and non-linear temporal associations between video clips, and 2) sparse associations of key information semantically related to the query. We propose a bidirectional decay retention mask to explicitly model temporal-distant context dependencies of video sequences, along with a learnable sparse retention mask to adaptively capture strong associations relevant to the target event. Extensive experiments conducted on five popular benchmarks ActivityNet Captions, TACoS, Charades-STA, ActivityNet Speech, and QVHighlights for MR tasks demonstrate the significant improvements achieved by our method over existing approaches. Code is available at https://github.com/xian-sh/MRNet.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- MOL-Mamba: Enhancing Molecular Representation with Structural & Electronic InsightsJingjing Hu, Dan Guo, Zhan Si, Deguang Liu 等AAAI 2025 · 被引用 9 次
- GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal GroundingRong Fan, Kaiyan Xiao, Minghao Zhu, Liuyi Wang 等CVPR 2026 · 被引用 1 次
- Object-Shot Enhanced Grounding Network for Egocentric VideoYisen Feng, Haoyu Zhang, Meng Liu, Weili Guan 等CVPR 2025
- Anchor-Aware Similarity Cohesion in Target Frames Enables Predicting Temporal Moment Boundaries in 2DJiawei Tan, Hongxing Wang, Junwu Weng, Jiaxin Li 等CVPR 2025
相关 Paper
- CDTR: Semantic Alignment for Video Moment Retrieval Using Concept Decomposition TransformerRan Ran, Jiwei Wei, Xiangyi Cai, Xiang Guan 等AAAI 2025 · 被引用 6 次
- UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionYe Liu, Siyuan Li, Yang Wu, Chang Wen Chen 等CVPR 2022 · 被引用 150 次
- Lightweight Relational Proposal Network with Dual-Branch Distillation for Video Moment RetrievalYujia Zhu, Hao Yang, Yibo Zhao, Chunjie Ma 等ACM MM 2025
- PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric NetworkMingyao Zhou, Hao Sun, Wei Xie, Ming Dong 等NeurIPS 2025 · 被引用 1 次
- Structured Multi-Level Interaction Network for Video Moment Localization via Language QueryHao Wang, Zheng-Jun Zha, Liang Li, Dong Liu 等CVPR 2021
