Lune

EMNLP2025顶会

AbsVis - Benchmarking How Humans and Vision-Language Models "See" Abstract Concepts in Images

Tarun Tater, Diego Frassinelli, Sabine Schulte im Walde

2025年份
2被引次数

摘要

concepts like mercy and peace often lack clear visual grounding, thus making it challenging to study how they are associated with images. To address this, we introduce AbsVisa dataset of 675 images annotated with 14, 175 concept-explanation pairs from humans and two Vision-Language Models (VLMs: Qwen and LLaVA), where each concept is supported by a textual explanation. We compare human and VLM attributions in terms of diversity, abstractness, and alignment, and find that humans attribute more varied concepts. AbsVis also includes 2, 680 human preference judgments evaluating the quality of a subset of these annotations, showing that overlapping concepts (attributed by both humans and VLMs) are most preferred. Explanations clarify and strengthen the perceived attributions, both from humans and VLMs. Finally, we show that VLMs can approximate human preferences and use them to fine-tune VLMs via Direct Preference Optimization (DPO), yielding improved alignments with preferred concept-explanation pairs.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 61601ed1-e405-4697-89cb-1d39a820b28d

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖