Maskable Retentive Network for Video Moment Retrieval
Jingjing Hu, Dan Guo, Kun Li, Zhan Si, Xun Yang, Meng Wang
Abstract
Video Moment Retrieval (MR) tasks involve predicting the moment described by a given natural language or spoken language query in an untrimmed video. In this paper, we propose a novel Maskable Retentive Network (MRNet) to address two key challenges in MR tasks: cross-modal guidance and video sequence modeling. Our approach introduces a new retention mechanism into the multimodal Transformer architecture, incorporating modality-specific attention modes. Specifically, we employ the Unlimited Attention for language-related attention regions to maximize cross-modal mutual guidance. Then, we introduce the Maskable Retention for video-only attention region to enhance video sequence modeling, that is, recognizing two crucial characteristics of video sequences: 1) bidirectional, decaying, and non-linear temporal associations between video clips, and 2) sparse associations of key information semantically related to the query. We propose a bidirectional decay retention mask to explicitly model temporal-distant context dependencies of video sequences, along with a learnable sparse retention mask to adaptively capture strong associations relevant to the target event. Extensive experiments conducted on five popular benchmarks ActivityNet Captions, TACoS, Charades-STA, ActivityNet Speech, and QVHighlights for MR tasks demonstrate the significant improvements achieved by our method over existing approaches. Code is available at https://github.com/xian-sh/MRNet.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- MOL-Mamba: Enhancing Molecular Representation with Structural & Electronic InsightsJingjing Hu, Dan Guo, Zhan Si, Deguang Liu et al.AAAI 2025 · 9 citations
- GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal GroundingRong Fan, Kaiyan Xiao, Minghao Zhu, Liuyi Wang et al.CVPR 2026 · 1 citation
- Object-Shot Enhanced Grounding Network for Egocentric VideoYisen Feng, Haoyu Zhang, Meng Liu, Weili Guan et al.CVPR 2025
- Anchor-Aware Similarity Cohesion in Target Frames Enables Predicting Temporal Moment Boundaries in 2DJiawei Tan, Hongxing Wang, Junwu Weng, Jiaxin Li et al.CVPR 2025
Related papers
- CDTR: Semantic Alignment for Video Moment Retrieval Using Concept Decomposition TransformerRan Ran, Jiwei Wei, Xiangyi Cai, Xiang Guan et al.AAAI 2025 · 6 citations
- UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionYe Liu, Siyuan Li, Yang Wu, Chang Wen Chen et al.CVPR 2022 · 150 citations
- Lightweight Relational Proposal Network with Dual-Branch Distillation for Video Moment RetrievalYujia Zhu, Hao Yang, Yibo Zhao, Chunjie Ma et al.ACM MM 2025
- PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric NetworkMingyao Zhou, Hao Sun, Wei Xie, Ming Dong et al.NeurIPS 2025 · 1 citation
- Structured Multi-Level Interaction Network for Video Moment Localization via Language QueryHao Wang, Zheng-Jun Zha, Liang Li, Dong Liu et al.CVPR 2021
