AdaCLIP: Towards Pragmatic Multimodal Video Retrieval
Zhiming Hu, Angela Ning Ye, Salar Hosseini Khorasgani, Iqbal Mohomed
Abstract
Incorporating large image-text foundation models such as CLIP has substantially improved the performance of the multimodal video retrieval task. However, how to practically sample the frames from a video and aggregate the frame features into a video representation is still an open research question. In particular, real-world deployment scenarios, such as embodiment within consumer electronics or cloud-based inference pipelines, require two key facets of retrieval (representation building and search) to be computationally light and fast. In this paper, we propose AdaCLIP, a computation- and latency-aware system for pragmatic multimodal video retrieval. AdaCLIP consists of a learning-based frame selection module to select informative frames and a query-independent frame aggregation module to obtain strong video representations from the frame features. Specifically, in the frame selection module, we introduce a differentiable Hard-Top-k algorithm to sample a subset of the frames while optimizing the performance of the video retrieval task in an end-to-end manner. Moreover, to be latency-aware, we also propose a query-independent lightweight approach, MLP-Score, to aggregate the frame features into the video representation, which offers up to 142x speedup on GPU and 822x speedup on CPU in similarity search time compared to query-dependent matching methods. Experimental results on several popular video retrieval datasets confirm the effectiveness of AdaCLIP.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6631eadc-66ad-439d-af3c-706d3e8d29e4Cited by top-tier papers3
- Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text RetrievalYang Du, Yuqi Liu, Qin JinACM MM 2024 · 4 citations
- CAPT: Confusion-Aware Prompt Tuning for Reducing Vision-Language MisalignmentMaoyuan Shao, Yutong Gao, Xinyang Huang, Lijuan Sun et al.CVPR 2026 · 1 citation
- Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level CaptionsChan Hur, Jeong-Hun Hong, Dong-hun Lee, Dabin Kang et al.CVPR 2025
Related papers
- Holistic Features are Almost Sufficient for Text-to-Video RetrievalKaibin Tian, Ruixiang Zhao, Zijie Xin, Bangxiang Lan et al.CVPR 2024 · 15 citations
- Q-Frame: Query-Aware Frame Selection and Multi-Resolution Adaptation for Video-LLMsShaojie Zhang, Jiahui Yang, Jianqin Yin, Zhenbo Luo et al.ICCV 2025 · 15 citations
- Efficient Frame Selection for Long Video Understanding via Reinforcement LearningYaxuan Qin, Hefei Li, Wenqi Mu, Yancheng HeCVPR 2026 · 6 citations
- Prompt Switch: Efficient CLIP Adaptation for Text-Video RetrievalChaorui Deng, Qi Chen, Pengda Qin, Da Chen et al.ICCV 2023 · 52 citations
- DGL: Dynamic Global-Local Prompt Tuning for Text-Video RetrievalXiangpeng Yang, Linchao Zhu, Xiaohan Wang, Yi YangAAAI 2024 · 53 citations
