CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
Shuai Zhao, Linchao Zhu, Xiaohan Wang, Yi Yang
摘要
Recently, large-scale pre-training methods like CLIP have made great progress in multi-modal research such as text-video retrieval. In CLIP, transformers are vital for modeling complex multi-modal relations. However, in the vision transformer of CLIP, the essential visual tokenization process, which produces discrete visual token sequences, generates many homogeneous tokens due to the redundancy nature of consecutive and similar frames in videos. This significantly increases computation costs and hinders the deployment of video retrieval models in web applications. In this paper, to reduce the number of redundant video tokens, we design a multi-segment token clustering algorithm to find the most representative tokens and drop the non-essential ones. As the frame redundancy occurs mostly in consecutive frames, we divide videos into multiple segments and conduct segment-level clustering. Center tokens from each segment are later concatenated into a new sequence, while their original spatial-temporal relations are well maintained. We instantiate two clustering algorithms to efficiently find deterministic medoids and iteratively partition groups in high dimensional space. Through this token clustering and center selection procedure, we successfully reduce computation costs by removing redundant visual tokens. This method further enhances segment-level semantic alignment between video and text representations, enforcing the spatio-temporal interactions of tokens from within-segment frames. Our method, coined as CenterCLIP, surpasses existing state-of-the-art by a large margin on typical text-video benchmarks, while reducing the training memory cost by 35% and accelerating the inference speed by 14% at the best case. The code is available at https://github.com/mzhaoshuai/CenterCLIP https://github.com/mzhaoshuai/CenterCLIP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper54
- Revisiting Classifier: Transferring Vision-Language Models for Video RecognitionWenhao Wu, Zhun Sun, Wanli OuyangAAAI 2023 · 被引用 141 次
- UATVR: Uncertainty-Adaptive Text-Video RetrievalBo Fang, Wenhao Wu, Chang Liu, Yu Zhou 等ICCV 2023 · 被引用 98 次
- Spectrum-guided Multi-granularity Referring Video Object SegmentationBo Miao, Mohammed Bennamoun, Yongsheng Gao, Ajmal MianICCV 2023 · 被引用 75 次
- Partially Relevant Video RetrievalJianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang 等ACM MM 2022 · 被引用 65 次
- Progressive Spatio-Temporal Prototype Matching for Text-Video RetrievalPandeng Li, Chen-Wei Xie, Liming Zhao, Hongtao Xie 等ICCV 2023 · 被引用 62 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
相关 Paper
- TempMe: Video Temporal Token Merging for Efficient Text-Video RetrievalLeqi Shen, Tianxiang Hao, Tao He, Sicheng Zhao 等ICLR 2025
- Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language PretrainingWeijun Zhuang, Yuqing Huang, Weikang Meng, Xin Li 等CVPR 2026 · 被引用 3 次
- SMAUG: Sparse Masked Autoencoder for Efficient Video-Language Pre-trainingYuanze Lin, Chen Wei, Huiyu Wang, Alan L. Yuille 等ICCV 2023 · 被引用 18 次
- Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low RetentionJunhao Du, Jialong Xue, Anqi Li, Jincheng Dai 等CVPR 2026 · 被引用 7 次
- VoP: Text-Video Co-Operative Prompt Tuning for Cross-Modal RetrievalSiteng Huang, Biao Gong, Yulin Pan, Jianwen Jiang 等CVPR 2023
