MUSE: Mamba Is Efficient Multi-scale Learner for Text-video Retrieval
Haoran Tang, Meng Cao, Jinfa Huang, Ruyang Liu, Peng Jin, Ge Li, Xiaodan Liang
Abstract
Text-Video Retrieval (TVR) aims to align and associate relevant video content with corresponding natural language queries. Most existing TVR methods are based on large-scale pre-trained vision-language models (e.g., CLIP). However, due to CLIP's inherent plain structure, few TVR methods explore the multi-scale representations which offer richer contextual information for a more thorough understanding. To this end, we propose MUSE, a multi-scale mamba with linear computational complexity for efficient cross-resolution modeling. Specifically, the multi-scale representations are generated by applying a feature pyramid on the last single-scale feature map. Then, we employ the Mamba structure as an efficient multi-scale learner to jointly learn scale-wise representations. Furthermore, we conduct comprehensive studies to investigate different model structures and designs. Extensive results on three popular benchmarks have validated the superiority of MUSE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7eeb4426-eec5-4775-a0a7-886cec36aa58Cited by top-tier papers7
- MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed VideosArkaprava Sinha, Monish Soundar Raj, Pu Wang, Ahmed Helmy et al.CVPR 2026 · 5 citations
- Video Spatial Reasoning with Object-Centric 3D RolloutHaoran Tang, Meng Cao, Ruyang Liu, Xiaoxi Liang et al.AAAI 2026 · 3 citations
- Temporal Calibrating and Distilling for Scene-Text Aware Text-Video RetrievalZhiqian Zhao, Liang Li, Lei Shen, Xichun Sheng et al.AAAI 2026 · 1 citation
- DHCM-CACL: Dynamic Hierarchical Cross-modal Mamba with Confidence-Adaptive Contrastive Learning for Multimodal Emotion RecognitionBaiqiang Wu, Yang LiAAAI 2026 · 1 citation
- MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video RetrievalReno Kriz, Kate Sanders, David Etter, Kenton Murray et al.CVPR 2025
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
Related papers
- CLIMB-ReID: A Hybrid CLIP-Mamba Framework for Person Re-IdentificationChenyang Yu, Xuehu Liu, Jiawen Zhu, Yuhao Wang et al.AAAI 2025 · 17 citations
- Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation LearningKaibin Tian, Yanhua Cheng, Yi Liu, Xinglin Hou et al.AAAI 2024 · 19 citations
- UATVR: Uncertainty-Adaptive Text-Video RetrievalBo Fang, Wenhao Wu, Chang Liu, Yu Zhou et al.ICCV 2023 · 98 citations
- MEME: Multi-Encoder Multi-Expert Framework with Data Augmentation for Video RetrievalSeong-Min Kang, Yoon-Sik ChoSIGIR 2023 · 7 citations
- Prompt Switch: Efficient CLIP Adaptation for Text-Video RetrievalChaorui Deng, Qi Chen, Pengda Qin, Da Chen et al.ICCV 2023 · 52 citations
