Visual Lexicon: Rich Image Features in Language Space
Xudong Wang, Xingyi Zhou, Alireza Fathi, Trevor Darrell, Cordelia Schmid
Abstract
Figure 1 . Given the cute corgi painting in the top left corner, how can we extract a visual representation that captures semantic-level information -such as object categories and layouts -while preserving rich visual details like image styles, textures and colors? We introduce ViLex model that generates image representations in the text vocabulary space, acting as a new visual "language", while retaining intricate visual details that are difficult, if not impossible, to convey in natural language. The set of images (generated under different diffusion noises) in the 2×2 grid, which are highly semantically and visually similar to each other, is created by using ViLex as "text" prompts for text-to-image diffusion models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Reconstruction Alignment Improves Unified Multimodal ModelsJi Xie, Trevor Darrell, Luke Zettlemoyer, XuDong WangICLR 2026 · 52 citations
- Unified Multimodal Models as Auto-EncodersZhiyuan Yan, Kaiqing Lin, Zongjian Li, Junyan Ye et al.CVPR 2026 · 12 citations
- CaTok: Taming Mean Flows for One-Dimensional Causal Image TokenizationYitong Chen, Zuxuan Wu, Xipeng Qiu, Yu-Gang JiangCVPR 2026 · 6 citations
- Vision‑Language‑Vision Auto‑Encoder: Scalable Knowledge Distillation from Diffusion ModelsTiezheng Zhang, Yitong Li, Yu-Cheng Chou, Jieneng Chen et al.NeurIPS 2025 · 5 citations
- GenIR: Generative Visual Feedback for Mental Image RetrievalDiji Yang, Minghao Liu, Chung-Hsiang Lo, Yi Zhang et al.NeurIPS 2025 · 4 citations
Builds on42
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Open-vocabulary Object Segmentation with Diffusion ModelsZiyi Li, Qinye Zhou, Xiaoyun Zhang, Ya Zhang et al.ICCV 2023 · 98 citations
- StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation LearnersYonglong Tian, Lijie Fan, Phillip Isola, Huiwen Chang et al.NeurIPS 2023 · 251 citations
- Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image ModelsGuy Kaplan, Michael Toker, Yuval Reif, Yonatan Belinkov et al.ACL 2026 · 4 citations
- Compositional Text-to-Image Generation with Dense Blob RepresentationsWeili Nie, Sifei Liu, Morteza Mardani, Chao Liu et al.ICML 2024 · 44 citations
- LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed PromptsHanan Gani, Shariq Farooq Bhat, Muzammal Naseer, Salman Khan et al.ICLR 2024 · 61 citations
