Visual Lexicon: Rich Image Features in Language Space
Xudong Wang, Xingyi Zhou, Alireza Fathi, Trevor Darrell, Cordelia Schmid
摘要
Figure 1 . Given the cute corgi painting in the top left corner, how can we extract a visual representation that captures semantic-level information -such as object categories and layouts -while preserving rich visual details like image styles, textures and colors? We introduce ViLex model that generates image representations in the text vocabulary space, acting as a new visual "language", while retaining intricate visual details that are difficult, if not impossible, to convey in natural language. The set of images (generated under different diffusion noises) in the 2×2 grid, which are highly semantically and visually similar to each other, is created by using ViLex as "text" prompts for text-to-image diffusion models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Reconstruction Alignment Improves Unified Multimodal ModelsJi Xie, Trevor Darrell, Luke Zettlemoyer, XuDong WangICLR 2026 · 被引用 52 次
- Unified Multimodal Models as Auto-EncodersZhiyuan Yan, Kaiqing Lin, Zongjian Li, Junyan Ye 等CVPR 2026 · 被引用 12 次
- CaTok: Taming Mean Flows for One-Dimensional Causal Image TokenizationYitong Chen, Zuxuan Wu, Xipeng Qiu, Yu-Gang JiangCVPR 2026 · 被引用 6 次
- Vision‑Language‑Vision Auto‑Encoder: Scalable Knowledge Distillation from Diffusion ModelsTiezheng Zhang, Yitong Li, Yu-Cheng Chou, Jieneng Chen 等NeurIPS 2025 · 被引用 5 次
- GenIR: Generative Visual Feedback for Mental Image RetrievalDiji Yang, Minghao Liu, Chung-Hsiang Lo, Yi Zhang 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper42
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Open-vocabulary Object Segmentation with Diffusion ModelsZiyi Li, Qinye Zhou, Xiaoyun Zhang, Ya Zhang 等ICCV 2023 · 被引用 98 次
- StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation LearnersYonglong Tian, Lijie Fan, Phillip Isola, Huiwen Chang 等NeurIPS 2023 · 被引用 251 次
- Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image ModelsGuy Kaplan, Michael Toker, Yuval Reif, Yonatan Belinkov 等ACL 2026 · 被引用 4 次
- Compositional Text-to-Image Generation with Dense Blob RepresentationsWeili Nie, Sifei Liu, Morteza Mardani, Chao Liu 等ICML 2024 · 被引用 44 次
- LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed PromptsHanan Gani, Shariq Farooq Bhat, Muzammal Naseer, Salman Khan 等ICLR 2024 · 被引用 61 次
