AbsVis - Benchmarking How Humans and Vision-Language Models "See" Abstract Concepts in Images
Tarun Tater, Diego Frassinelli, Sabine Schulte im Walde
摘要
concepts like mercy and peace often lack clear visual grounding, thus making it challenging to study how they are associated with images. To address this, we introduce AbsVisa dataset of 675 images annotated with 14, 175 concept-explanation pairs from humans and two Vision-Language Models (VLMs: Qwen and LLaVA), where each concept is supported by a textual explanation. We compare human and VLM attributions in terms of diversity, abstractness, and alignment, and find that humans attribute more varied concepts. AbsVis also includes 2, 680 human preference judgments evaluating the quality of a subset of these annotations, showing that overlapping concepts (attributed by both humans and VLMs) are most preferred. Explanations clarify and strengthen the perceived attributions, both from humans and VLMs. Finally, we show that VLMs can approximate human preferences and use them to fine-tune VLMs via Direct Preference Optimization (DPO), yielding improved alignments with preferred concept-explanation pairs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- Improve Vision Language Model Chain-of-thought ReasoningRuohong Zhang, Bowen Zhang, Yanghao Li, Haotian Zhang 等ACL 2025 · 被引用 135 次
相关 Paper
- Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across ModalitiesYiting Qu, Michael Backes, Yang ZhangUSENIX Security 2025
- Self-Supervised Visual Preference AlignmentKe Zhu, Liang Zhao, Zheng Ge, Xiangyu ZhangACM MM 2024 · 被引用 7 次
- Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language ModelsJeonghwan Kim, Heng JiEMNLP 2024 · 被引用 4 次
- OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human PreferenceXiangyu Zhao, Shengyuan Ding, Zicheng Zhang, Haian Huang 等ACL 2025 · 被引用 24 次
- LLaVA-Ultra: Large Chinese Language and Vision Assistant for UltrasoundXuechen Guo, Wenhao Chai, Shiyan Li, Gaoang WangACM MM 2024 · 被引用 18 次
