AbsVis - Benchmarking How Humans and Vision-Language Models "See" Abstract Concepts in Images
Tarun Tater, Diego Frassinelli, Sabine Schulte im Walde
Abstract
concepts like mercy and peace often lack clear visual grounding, thus making it challenging to study how they are associated with images. To address this, we introduce AbsVisa dataset of 675 images annotated with 14, 175 concept-explanation pairs from humans and two Vision-Language Models (VLMs: Qwen and LLaVA), where each concept is supported by a textual explanation. We compare human and VLM attributions in terms of diversity, abstractness, and alignment, and find that humans attribute more varied concepts. AbsVis also includes 2, 680 human preference judgments evaluating the quality of a subset of these annotations, showing that overlapping concepts (attributed by both humans and VLMs) are most preferred. Explanations clarify and strengthen the perceived attributions, both from humans and VLMs. Finally, we show that VLMs can approximate human preferences and use them to fine-tune VLMs via Direct Preference Optimization (DPO), yielding improved alignments with preferred concept-explanation pairs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61601ed1-e405-4697-89cb-1d39a820b28dBuilds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- Improve Vision Language Model Chain-of-thought ReasoningRuohong Zhang, Bowen Zhang, Yanghao Li, Haotian Zhang et al.ACL 2025 · 135 citations
Related papers
- Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across ModalitiesYiting Qu, Michael Backes, Yang ZhangUSENIX Security 2025
- Self-Supervised Visual Preference AlignmentKe Zhu, Liang Zhao, Zheng Ge, Xiangyu ZhangACM MM 2024 · 7 citations
- Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language ModelsJeonghwan Kim, Heng JiEMNLP 2024 · 4 citations
- OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human PreferenceXiangyu Zhao, Shengyuan Ding, Zicheng Zhang, Haian Huang et al.ACL 2025 · 24 citations
- LLaVA-Ultra: Large Chinese Language and Vision Assistant for UltrasoundXuechen Guo, Wenhao Chai, Shiyan Li, Gaoang WangACM MM 2024 · 18 citations
