MS-DETR: Natural Language Video Localization with Sampling Moment-Moment Interaction
Jing Wang, Aixin Sun, Hao Zhang, Xiaoli Li
Abstract
Given a query, the task of Natural Language Video Localization (NLVL) is to localize a temporal moment in an untrimmed video that semantically matches the query. In this paper, we adopt a proposal-based solution that generates proposals (i.e., candidate moments) and then select the best matching proposal. On top of modeling the cross-modal interaction between candidate moments and the query, our proposed Moment Sampling DETR (MS-DETR) enables efficient moment-moment relation modeling. The core idea is to sample a subset of moments guided by the learnable templates with an adopted DETR (DEtection TRansformer) framework. To achieve this, we design a multiscale visual-linguistic encoder, and an anchorguided moment decoder paired with a set of learnable templates. Experimental results on three public datasets demonstrate the superior performance of MS-DETR. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9745c80-65eb-4c8c-8061-14d13c3b24d7Cited by top-tier papers5
- Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment RetrievalZhihang Liu, Jun Li, Hongtao Xie, Pandeng Li et al.AAAI 2024 · 49 citations
- Generative Video Diffusion for Unseen Novel Semantic Video Moment RetrievalDezhao Luo, Shaogang Gong, Jiabo Huang, Hailin Jin et al.AAAI 2025 · 4 citations
- Short Video Segment-level User Dynamic Interests Modeling in Personalized RecommendationZhiyu He, Zhixin Ling, Jiayu Li, Zhiqiang Guo et al.SIGIR 2025 · 4 citations
- Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment RetrievalJunan Lin, Daizong Liu, Xianke Chen, Xiaoye Qu et al.ACM MM 2025 · 3 citations
- Exploiting Intrinsic Multilateral Logical Rules for Weakly Supervised Natural Language Video LocalizationZhe Xu, Kun Wei, Xu Yang, Cheng DengACL 2024
Builds on16
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
Related papers
- Natural Language Video Localization with Learnable Moment ProposalsShaoning Xiao, Long Chen, Jian Shao, Yueting Zhuang et al.EMNLP 2021 · 44 citations
- TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight DetectionHao Sun, Mingyao Zhou, Wenjing Chen, Wei XieAAAI 2024
- Reducing Intrinsic and Extrinsic Data Biases for Moment Localization with Natural LanguageJiong Yin, Liang Li, Jiehua Zhang, Chenggang Yan et al.ACM MM 2023 · 4 citations
- MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic LearningHongxu Ma, Guanshuo Wang, Fufu Yu, Qiong Jia et al.ACM MM 2025 · 9 citations
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 279 citations
