Bridging the gap to real-world language-grounded visual concept learning
Whie Jung, Semin Kim, Junee Kim, Seunghoon Hong
摘要
Human intelligence effortlessly interprets visual scenes along a rich spectrum of semantic dimensions. However, existing approaches to language-grounded visual concept learning are limited to a few predefined primitive axes, such as color and shape, and are typically explored in synthetic datasets. In this work, we propose a scalable framework that adaptively identifies image-related concept axes and grounds visual concepts along these axes in real-world scenes. Leveraging a pretrained vision-language model and our universal prompting strategy, our framework identifies a diverse image-related axes without any prior knowledge. Our universal concept encoder adaptively binds visual features to the discovered axes without introducing additional model parameters for each concept. To ground visual concepts along the discovered axes, we optimize a compositional anchoring objective, which ensures that each axis can be independently manipulated without affecting others. We demonstrate the effectiveness of our framework on subsets of ImageNet, CelebA-HQ, and AFHQ, showcasing superior editing capabilities across diverse real-world concepts that are too varied to be manually predefined. Our method also exhibits strong compositional generalization, outperforming existing visual concept learning and text-based editing methods. The code is available at https://github.com/whieya/Language-grounded-VCL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- Language-Informed Visual Concept LearningSharon Lee, Yunzhi Zhang, Shangzhe Wu, Jiajun WuICLR 2024 · 被引用 13 次
- Advancing Textual Prompt Learning with Anchored AttributesZheng Li, Yibing Song, Ming-Ming Cheng, Xiang Li 等ICCV 2025 · 被引用 8 次
- Exploring Interpretability for Visual Prompt Tuning with Cross-layer ConceptsYubin Wang, Xinyang Jiang, De Cheng, Xiangqian Zhao 等ICLR 2026 · 被引用 1 次
- Decomposition of Concept-Level Rules in Visual ScenesFan Shi, Yuxuan Liang, Xiaolei Chen, Haiyang Yu 等ICLR 2026
- LaViP: Language-Grounded Visual PromptingNilakshan Kunananthaseelan, Jing Zhang, Mehrtash HarandiAAAI 2024 · 被引用 6 次
