Seeing the Abstract: Translating the Abstract Language for Vision Language Models
Davide Talon, Federico Girella, Ziyue Liu, Marco Cristani, Yiming Wang
Abstract
Natural language goes beyond dryly describing visual content. It contains rich abstract concepts to express feeling, creativity and properties that cannot be directly perceived. Yet, current research in Vision Language Models (VLMs) has not shed light on abstract-oriented language. Our research breaks new ground by uncovering its wide presence and under-estimated value, with extensive analysis. Particularly, we focus our investigation on the fashion domain, a highly-representative field with abstract expressions. By analyzing recent large-scale multimodal fashion datasets, we find that abstract terms have a dominant presence, rivaling the concrete ones, providing novel information, and being useful in the retrieval task. However, a critical challenge emerges: current general-purpose or fashion-specific VLMs are pre-trained with databases that lack sufficient abstract words in their text corpora, thus hindering their ability to effectively represent abstract-oriented language. We propose a training-free and model-agnostic method, Abstractto-Concrete Translator (ACT), to shift abstract representations towards well-represented concrete ones in the VLM latent space, using pre-trained models and existing multimodal databases. On the text-to-image retrieval task, despite being training-free, ACT outperforms the fine-tuned VLMs in both same-and cross-dataset settings, exhibiting its effectiveness with a strong generalization capability. Moreover, the improvement introduced by ACT is consistent with various VLMs, making it a plug-and-play solution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c1e2971-b670-4012-af29-e8c7f48bb7faCited by top-tier papers4
- LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text PairingFederico Girella, Davide Talon, Ziyue Liu, Zanxi Ruan et al.ICCV 2025 · 2 citations
- StructXLIP: Enhancing Vision-language Models with Multimodal Structural CuesZanxi Ruan, Songqun Gao, Qiuyu Kong, Yiming Wang et al.CVPR 2026 · 1 citation
- How to Take a Memorable Picture? Empowering Users with Actionable FeedbackFrancesco Laiti, Davide Talon, Jacopo Staiano, Elisa RicciCVPR 2026
- Visual Grounding for Object QuestionsMartin Nicolas Everaert, Xiruo Liu, Hiroyuki Takeda, Raja Bala et al.CVPR 2026
Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung et al.NeurIPS 2022 · 834 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- Effective conditioned and composed image retrieval combining CLIP-based featuresAlberto Baldrati, Marco Bertini, Tiberio Uricchio, Alberto Del BimboCVPR 2022 · 139 citations
Related papers
- Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation AdapterWeizhi Zhong, Huan Yang, Zheng Liu, Huiguo He et al.ICLR 2026 · 17 citations
- FashionVLP: Vision Language Transformer for Fashion Retrieval with FeedbackSonam Goenka, Zhaoheng Zheng, Ayush Jaiswal, Rakesh Chada et al.CVPR 2022 · 88 citations
- AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationYuhan Zhu, Yuyang Ji, Zhiyu Zhao, Gangshan Wu et al.NeurIPS 2024 · 45 citations
- MoTrans: Customized Motion Transfer with Text-driven Video Diffusion ModelsXiaomin Li, Xu Jia, Qinghe Wang, Haiwen Diao et al.ACM MM 2024 · 6 citations
- LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsMuhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Horst Possegger et al.NeurIPS 2023 · 63 citations
