Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
Shizhan Gong, Yankai Jiang, Qi Dou, Farzan Farnia
摘要
Vision-language models, such as CLIP, have achieved significant success in aligning visual and textual representations, becoming essential components of many multi-modal large language models (MLLMs) like LLaVA and Open-Flamingo. However, numerous studies have identified CLIP's limited fine-grained perception as a critical drawback, leading to substantial failures in downstream MLLMs. In contrast, visioncentric foundation models like DINOv2 demonstrate remarkable capabilities in capturing fine details from images. In this work, we propose a novel kernel-based method to align CLIP's visual representation with that of DINOv2, ensuring that the resulting embeddings maintain compatibility with text embeddings while enhancing perceptual capabilities. Our alignment objective is designed for efficient stochastic optimization. Following this image-only alignment fine-tuning, the visual encoder retains compatibility with the frozen text encoder and exhibits significant improvements in zero-shot object recognition, finegrained spatial reasoning, and localization. By integrating the aligned visual encoder, downstream MLLMs also demonstrate enhanced performance. The code and models are available at https: //github.com/peterant330/KUEA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Scendi Score: Prompt-Aware Diversity Evaluation Via Schur Complement of Clip EmbeddingsAzim Ospanov, Mohammad Jalali, Farzan FarniaICCV 2025 · 被引用 17 次
- SPARKE: Scalable Prompt-Aware Diversity and Novelty Guidance in Diffusion Models via RKE ScoreMohammad Jalali, Haoyu Lei, Amin Gohari, Farzan FarniaNeurIPS 2025 · 被引用 15 次
- When Kernels Multiply, Clusters Unify: Fusing Embeddings with the Kronecker ProductYouqi Wu, Jingwei Zhang, Farzan FarniaNeurIPS 2025 · 被引用 7 次
- DAK-UCB: Diversity-Aware Prompt Routing for LLMs and Generative ModelsDonya Jafari, Farzan FarniaICLR 2026 · 被引用 5 次
- MMD Guidance: Training-Free Distribution Adaptation for Diffusion Models via Maximum Mean Discrepancy GuidanceMatina Mahdizadeh Sani, Nima Jamali, Mohammad Jalali, Farzan FarniaICML 2026 · 被引用 4 次
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
相关 Paper
- Harnessing Frozen Unimodal Encoders for Flexible Multimodal AlignmentMayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Sanath Narayan 等CVPR 2025
- Do Vision and Language Encoders Represent the World Similarly?Mayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Mohamed El Amine Seddik 等CVPR 2024 · 被引用 3 次
- Unified Lexical Representation for Interpretable Visual-Language AlignmentYifan Li, Yikai Wang, Yanwei Fu, Dongyu Ru 等NeurIPS 2024 · 被引用 9 次
- DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language AlignmentCijo Jose, Théo Moutakanni, Dahyun Kang, Federico Baldassarre 等CVPR 2025
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang 等ICML 2024 · 被引用 7 次
