MS-DETR: Natural Language Video Localization with Sampling Moment-Moment Interaction
Jing Wang, Aixin Sun, Hao Zhang, Xiaoli Li
摘要
Given a query, the task of Natural Language Video Localization (NLVL) is to localize a temporal moment in an untrimmed video that semantically matches the query. In this paper, we adopt a proposal-based solution that generates proposals (i.e., candidate moments) and then select the best matching proposal. On top of modeling the cross-modal interaction between candidate moments and the query, our proposed Moment Sampling DETR (MS-DETR) enables efficient moment-moment relation modeling. The core idea is to sample a subset of moments guided by the learnable templates with an adopted DETR (DEtection TRansformer) framework. To achieve this, we design a multiscale visual-linguistic encoder, and an anchorguided moment decoder paired with a set of learnable templates. Experimental results on three public datasets demonstrate the superior performance of MS-DETR. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment RetrievalZhihang Liu, Jun Li, Hongtao Xie, Pandeng Li 等AAAI 2024 · 被引用 49 次
- Generative Video Diffusion for Unseen Novel Semantic Video Moment RetrievalDezhao Luo, Shaogang Gong, Jiabo Huang, Hailin Jin 等AAAI 2025 · 被引用 4 次
- Short Video Segment-level User Dynamic Interests Modeling in Personalized RecommendationZhiyu He, Zhixin Ling, Jiayu Li, Zhiqiang Guo 等SIGIR 2025 · 被引用 4 次
- Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment RetrievalJunan Lin, Daizong Liu, Xianke Chen, Xiaoye Qu 等ACM MM 2025 · 被引用 3 次
- Exploiting Intrinsic Multilateral Logical Rules for Weakly Supervised Natural Language Video LocalizationZhe Xu, Kun Wei, Xu Yang, Cheng DengACL 2024
它引用的顶会 Paper16
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng 等ICCV 2021 · 被引用 974 次
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo 等CVPR 2022 · 被引用 879 次
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
相关 Paper
- Natural Language Video Localization with Learnable Moment ProposalsShaoning Xiao, Long Chen, Jian Shao, Yueting Zhuang 等EMNLP 2021 · 被引用 44 次
- TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight DetectionHao Sun, Mingyao Zhou, Wenjing Chen, Wei XieAAAI 2024
- Reducing Intrinsic and Extrinsic Data Biases for Moment Localization with Natural LanguageJiong Yin, Liang Li, Jiehua Zhang, Chenggang Yan 等ACM MM 2023 · 被引用 4 次
- MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic LearningHongxu Ma, Guanshuo Wang, Fufu Yu, Qiong Jia 等ACM MM 2025 · 被引用 9 次
- Span-based Localizing Network for Natural Language Video LocalizationHao Zhang, Aixin Sun, Wei Jing, Joey Tianyi ZhouACL 2020 · 被引用 279 次
