Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
Xin Xiao, Bohong Wu, Jiacong Wang, Chunyuan Li, Xun Zhou, Haoyuan Guo
摘要
Existing image-text modality alignment in Vision Language Models (VLMs) treats each text token equally in an autoregressive manner. Despite being simple and effective, this method results in sub-optimal cross-modal alignment by over-emphasizing the text tokens that are less correlated with or even contradictory with the input images. In this paper, we advocate for assigning distinct contributions for each text token based on its visual correlation. Specifically, we present by contrasting image inputs, the difference in prediction logits on each text token provides strong guidance of visual correlation. We therefore introduce Contrastive ALignment (CAL), a simple yet effective re-weighting strategy that prioritizes training visually correlated tokens. Our experimental results demonstrate that CAL consistently improves different types of VLMs across different resolutions and model sizes on various benchmark datasets. Importantly, our method incurs minimal additional computational overhead, rendering it highly efficient compared to alternative data scaling strategies. Codes are available at https://github.com/foundation-multimodal-models/CAL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- VGR: Visual Grounded ReasoningJiacong Wang, Zijian Kang, Haochen Wang, Xiao Liang 等ICLR 2026 · 被引用 64 次
- TG-LLaVA: Text Guided LLaVA via Learnable Latent EmbeddingsDawei Yan, Pengcheng Li, Yang Li, Hao Chen 等AAAI 2025 · 被引用 10 次
- ProxyThinker: Test-Time Guidance through Small Visual ReasonersZilin Xiao, Jaywon Koo, Siru Ouyang, Jefferson Hernandez 等ICLR 2026 · 被引用 8 次
- Knowledge Transfer from Interaction LearningYilin Gao, Kangyi Chen, Zhongxing Peng, Hengjie Lu 等ICCV 2025 · 被引用 3 次
- ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language ModelsJunzhe Chen, Tianshu Zhang, Shiyu Huang, Yuwei Niu 等CVPR 2025
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document UnderstandingAhmed Masry, Juan A. Rodríguez, Tianyu Zhang, Suyuchen Wang 等NeurIPS 2025 · 被引用 7 次
- Token Embeddings Alignment for Cross-Modal RetrievalChen-Wei Xie, Jianmin Wu, Yun Zheng, Pan Pan 等ACM MM 2022 · 被引用 18 次
- FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal UnderstandingZheng Liu, Mengjie Liu, Jingzhou Chen, Jingwei Xu 等ICLR 2026 · 被引用 3 次
- Seeing What You Miss: Vision-Language Pre-training with Semantic Completion LearningYatai Ji, Rongcheng Tu, Jie Jiang, Weijie Kong 等CVPR 2023
- MODIX: A Training-Free Multimodal Information-Driven Positional Index Scaling for Vision-Language ModelsRuoxiang Huang, Zhen YuanCVPR 2026 · 被引用 2 次
