VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
Jiapeng Wang, Chengyu Wang, Kunzhe Huang, Jun Huang, Lianwen Jin
摘要
Contrastive Language-Image Pre-training (CLIP) has been widely studied and applied in numerous applications. However, the emphasis on brief summary texts during pre-training prevents CLIP from understanding long descriptions. This issue is particularly acute regarding videos given that videos often contain abundant detailed contents. In this paper, we propose the VideoCLIP-XL (eXtra Length) model, which aims to unleash the long-description understanding capability of video CLIP models. Firstly, we establish an automatic data collection system and gather a large-scale VILD pre-training dataset 1 with VIdeo and Long-Description pairs. Then, we propose Text-similarity-guided Primary Component Matching (TPCM) to better learn the distribution of feature space while expanding the long description capability. We also introduce two new tasks namely Detail-aware Description Ranking (DDR) and Hallucination-aware Description Ranking (HDR) for further understanding improvement. Finally, we construct a Long Video Description Ranking (LVDR) benchmark 2 for evaluating the long-description capability more comprehensively. Extensive experimental results on widely-used text-video retrieval benchmarks with both short and long descriptions and our LVDR benchmark can fully demonstrate the effectiveness of our method. 3 * Contribution during internship at Alibaba Cloud Computing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing AssessmentYinan Chen, Jiangning Zhang, Teng Hu, Yuxiang Zeng 等ICLR 2026 · 被引用 29 次
- PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware MechanismsYifei Xia, Shuchen Weng, Siqi Yang, Jingqi Liu 等NeurIPS 2025 · 被引用 24 次
- Dexterous World ModelsByungjun Kim, Taeksoo Kim, Junyoung Lee, Hanbyul JooCVPR 2026 · 被引用 17 次
- Audio-Sync Video Generation with Multi-Stream Temporal ControlShuchen Weng, Haojie Zheng, Zheng Chang, Si Li 等NeurIPS 2025 · 被引用 14 次
- T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video GenerationZhe Cao, Tao Wang, Jiaming Wang, Yanghai Wang 等ICML 2026 · 被引用 13 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
相关 Paper
- RWKV-CLIP: A Robust Vision-Language Representation LearnerTiancheng Gu, Kaicheng Yang, Xiang An, Ziyong Feng 等EMNLP 2024 · 被引用 11 次
- PixCLIP: Towards Fine-grained Vision-Language Understanding via Any-granularity Pixel-Text AlignmentYicheng Xiao, Yu Chen, Hao-Xuan Ma, Jiale Hong 等ICML 2026 · 被引用 4 次
- FG-CLIP: Fine-Grained Visual and Textual AlignmentChunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li 等ICML 2025
- OneLIP: Unlocking and Improving Long-Text Representations of CLIP via One-Stage AdaptationRenjie Pan, Jiayan Song, Hua YangAAAI 2026
- Enhanced Motion-Text Alignment for Image-to-Video Transfer LearningWei Zhang, Chaoqun Wan, Tongliang Liu, Xinmei Tian 等CVPR 2024 · 被引用 8 次
