Lune

IEEE VIS2025顶会

EncQA: Benchmarking Vision-Language Models on Visual Encodings for Charts

Kushin Mukherjee, Donghao Ren, Dominik Moritz, Yannick Assogba

2025年份
3被引次数
1顶会引用

摘要

What is the value of Var at D? Which bar is longest? Which circles have a value greater than X? What is the average value of Var? Which color is an outlier in terms of number of observations? Are var1 and var2 correlated? Retrieve Value Position Length Area Color (Quant.) Color (Nominal) Shape Find Extrema Filter Values Derived Value Find Anomaly Correlate Values

Fig. 1: ENCQA charts vary across tasks and encodings. Associated questions focus on the visual mapping used to encode the data, for example in the FIND ANOMALY task with the AREA encoding: "Which circle is an outlier relative to the rest in terms of area?"

Abstract-Multimodal vision-language models (VLMs) continue to achieve ever-improving scores on chart understanding benchmarks. Yet, we find that this progress does not fully capture the breadth of visual reasoning capabilities essential for interpreting charts. We introduce ENCQA, a novel benchmark informed by the visualization literature, designed to provide systematic coverage of visual encodings and analytic tasks that are crucial for chart understanding. ENCQA provides 2,076 synthetic question-answer pairs, enabling balanced coverage of six visual encoding channels (position, length, area, color quantitative, color nominal, and shape) and eight tasks (find extrema, retrieve value, find anomaly, filter values, compute derived value exact, compute derived value relative, correlate values, and correlate values relative). Our evaluation of 9 state-of-the-art VLMs reveals that performance varies significantly across encodings within the same task, as well as across tasks. Contrary to expectations, we observe that performance does not improve with model size for many task-encoding pairs. Our results suggest that advancing chart understanding requires targeted strategies addressing specific visual reasoning gaps, rather than solely scaling up model or dataset size.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper15

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖