A Vision Check-up for Language Models
Pratyusha Sharma, Tamar Rott Shaham, Manel Baradad, Adrián Rodríguez-Muñoz, Shivam Duggal, Phillip Isola, Antonio Torralba, Stephanie Fu
摘要
What does learning to model relationships between strings teach Large Language Models (LLMs) about the visual world? We systematically evaluate LLMs' abilities to generate and recognize an assortment of visual concepts of increasing complexity and then demonstrate how a preliminary visual representation learning system can be trained using models of text. As language models lack the ability to consume or output visual information as pixels, we use code to represent images in our study. Although LLM-generated images do not look like natural images, results on image generation and the ability of models to correct these generated images indicate that precise modeling of strings can teach language models about numerous aspects of the visual world. Furthermore, experiments on self-supervised visual representation learning, utilizing images generated with text models, highlight the potential to train vision models capable of making semantic assessments of natural images using just LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Vision Language Models are BiasedAn Vo, Khai-Nguyen Nguyen, Mohammad Reza Taesiri, Thi Tuong Vy Dang 等ICLR 2026 · 被引用 68 次
- Perception-Guided Jailbreak Against Text-to-Image ModelsYihao Huang, Le Liang, Tianlin Li, Xiaojun Jia 等AAAI 2025 · 被引用 34 次
- UIClip: A Data-driven Model for Assessing User Interface DesignJason Wu, Yi-Hao Peng, Xin Yue Amanda Li, Amanda Swearngin 等UIST 2024 · 被引用 29 次
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-trainingJunlin Han, Shengbang Tong, David Fan, Yufan Ren 等ICLR 2026 · 被引用 25 次
- Tikzero: Zero-Shot Text-Guided Graphics Program SynthesisJonas Belouadi, Eddy Ilg, Margret Keuper, Hideki Tanaka 等ICCV 2025 · 被引用 24 次
它引用的顶会 Paper13
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami 等NeurIPS 2021 · 被引用 1,020 次
- StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation LearnersYonglong Tian, Lijie Fan, Phillip Isola, Huiwen Chang 等NeurIPS 2023 · 被引用 251 次
相关 Paper
- Mapping Language Models to Grounded Conceptual SpacesRoma Patel, Ellie PavlickICLR 2022 · 被引用 197 次
- Elucidating the design space of language models for image generationXuantong Liu, Shaozhe Hao, Xianbiao Qi, Tianyang Hu 等ICML 2025
- Contrastive Localized Language-Image Pre-TrainingHong-You Chen, Zhengfeng Lai, Haotian Zhang, Xinze Wang 等ICML 2025
- Learning Visual Representations via Language-Guided SamplingMohamed El Banani, Karan Desai, Justin JohnsonCVPR 2023
- On LLMs’ Internal Representation of Code CorrectnessFrancisco Ribeiro, Claudio Spiess, Premkumar Devanbu, Sarah NadiICSE 2026
