FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
Bingchao Wang, Zhiwei Ning, Jianyu Ding, Xuanang Gao, Yin Li, Dongsheng Jiang, Jie Yang, Wei Liu
摘要
CLIP has shown promising performance across many shorttext tasks in a zero-shot manner. However, limited by the input length of the text encoder, CLIP struggles on understream tasks with long-text inputs ( tokens). To remedy this issue, we propose FIX-CLIP, which includes three novel modules: (1) A dual-branch training pipeline that aligns short and long texts with masked and raw images, respectively, which boosts the long-text representation while preserving the short-text ability. (2) Multiple learnable regional prompts with unidirectional masks in Transformer layers for regional information extraction. (3) A hierarchical feature alignment module in the intermediate encoder layers to promote the consistency of multi-scale features. Furthermore, we collect 30M images and utilize existing MLLMs to synthesize long-text captions for training. Extensive experiments show that FIX-CLIP achieves state-of-the-art performance on both long-text and short-text retrieval benchmarks. For downstream applications, we reveal that FIX-CLIP's text encoder delivers promising performance in a plug-andplay manner for diffusion models with long-text input. The code is available at https://github.com/bcwang-sjtu/Fix-CLIP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable SpecializationAnzhe Cheng, Shukai Duan, Shixuan Li, Chenzhong Yin 等CVPR 2026 · 被引用 8 次
- DynamicID: Zero-Shot Multi-ID Image Personalization With Flexible Facial EditabilityXirui Hu, Jiahao Wang, Hao Chen, Weizhan Zhang 等ICCV 2025 · 被引用 3 次
- Hybrid Layout Control for Diffusion Transformer: Fewer Annotations, Superior AestheticsKeming Wu, Junwen Chen, Zhanhao Liang, Yinuo Wang 等ICCV 2025 · 被引用 2 次
- PowerCLIP: Powerset Alignment for Contrastive Pre-TrainingMasaki Kawamura, Nakamasa Inoue, Rintaro Yanagi, Hirokatsu Kataoka 等CVPR 2026 · 被引用 1 次
- CLIP Is Shortsighted: Paying Attention Beyond the First SentenceMarc-Antoine Lavoie, Anas Mahmoud, Aldo Zaimi, Arsene Fansi Tchango 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- PixCLIP: Towards Fine-grained Vision-Language Understanding via Any-granularity Pixel-Text AlignmentYicheng Xiao, Yu Chen, Hao-Xuan Ma, Jiale Hong 等ICML 2026 · 被引用 4 次
- Retaining Knowledge and Enhancing Long-Text Representations in CLIP through Dual-Teacher DistillationYuheng Feng, Changsong Wen, Zelin Peng, Li jiaye 等CVPR 2025
- TULIP: Token-length Upgraded CLIPIvona Najdenkoska, Mohammad Mahdi Derakhshani, Yuki M. Asano, Nanne van Noord 等ICLR 2025
- FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text InputsMothilal Asokan, Kebin Wu, Fatima AlbreikiCVPR 2025
- OneLIP: Unlocking and Improving Long-Text Representations of CLIP via One-Stage AdaptationRenjie Pan, Jiayan Song, Hua YangAAAI 2026
