Language-Informed Visual Concept Learning
Sharon Lee, Yunzhi Zhang, Shangzhe Wu, Jiajun Wu
Abstract
Our understanding of the visual world is centered around various concept axes, characterizing different aspects of visual entities. While different concept axes can be easily specified by language, e.g., color, the exact visual nuances along each axis often exceed the limitations of linguistic articulations, e.g., a particular style of painting. In this work, our goal is to learn a language-informed visual concept representation, by simply distilling large pre-trained vision-language models. Specifically, we train a set of concept encoders to encode the information pertinent to a set of language-informed concept axes, with an objective of reproducing the input image through a pre-trained Text-to-Image (T2I) model. To encourage better disentanglement of different concept encoders, we anchor the concept embeddings to a set of text embeddings obtained from a pre-trained Visual Question Answering (VQA) model. At inference time, the model extracts concept embeddings along various axes from new test images, which can be remixed to generate images with novel compositions of visual concepts. With a lightweight test-time finetuning procedure, it can also generalize to novel concepts unseen at training. Project page at https://cs.stanford.edu/ ˜yzzhang/projects/concept-axes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef5be4f8-42aa-4d20-aed2-4132ac0fd05aCited by top-tier papers9
- Compositional Image Decomposition with Diffusion ModelsJocelin Su, Nan Liu, Yanbo Wang, Joshua B. Tenenbaum et al.ICML 2024 · 16 citations
- IP-Composer: Semantic Composition of Visual ConceptsSara Dorfman, Dana Cohen-Bar, Rinon Gal, Daniel Cohen-OrSIGGRAPH 2025 · 4 citations
- pOps: Photo-Inspired Diffusion OperatorsElad Richardson, Yuval Alaluf, Ali Mahdavi-Amiri, Daniel Cohen-OrSIGGRAPH 2025 · 1 citation
- Inspiration Seeds: Learning Non-Literal Visual Combinations for Generative ExplorationKfir Goldberg, Elad Richardson, Yael VinkerSIGGRAPH 2026 · 1 citation
- PostEdit: Posterior Sampling for Efficient Zero-Shot Image EditingFeng Tian, Yixuan Li, Yichao Yan, Shanyan Guan et al.ICLR 2025
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Bridging the gap to real-world language-grounded visual concept learningWhie Jung, Semin Kim, Junee Kim, Seunghoon HongNeurIPS 2025
- Overcoming Language Priors in VQA via Decomposed Linguistic RepresentationsChenchen Jing, Yuwei Wu, Xiaoxun Zhang, Yunde Jia et al.AAAI 2020 · 115 citations
- Dual-Process Image GenerationGrace Luo, Jonathan Granskog, Aleksander Holynski, Trevor DarrellICCV 2025 · 2 citations
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik et al.ICLR 2023 · 464 citations
- Improving Personalized Search with Regularized Low-Rank Parameter UpdatesFiona Ryan, Josef Sivic, Fabian Caba Heilbron, Judy Hoffman et al.CVPR 2025
