VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations
Yupeng Xie, Zhiyang Zhang, Yifan Wu, Sirong Lu, Jiayi Zhang, Zhaoyang Yu, Jinlin Wang, Sirui Hong, Bang Liu, Chenglin Wu, Yuyu Luo
Abstract
Visualization, a domain-specific yet widely used form of imagery, is an effective way to turn complex datasets into intuitive insights, and its value depends on whether data are faithfully represented, clearly communicated, and aesthetically designed. However, evaluating visualization quality is challenging: unlike natural images, it requires simultaneous judgment across data encoding accuracy, information expressiveness, and visual aesthetics. Although multimodal large language models (MLLMs) have shown promising performance in aesthetic assessment of natural images, no systematic benchmark exists for measuring their capabilities in evaluating visualizations. To address this, we propose VISJUDGE-BENCH, the first comprehensive benchmark for evaluating MLLMs' performance in assessing visualization aesthetics and quality. It contains 3,090 expert-annotated samples from real-world scenarios, covering single visualizations, multiple visualizations, and dashboards across 32 chart types. Systematic testing on this benchmark reveals that even the most advanced MLLMs (such as GPT-5) still exhibit significant gaps compared to human experts in judgment, with a Mean Absolute Error (MAE) of 0.553 and a correlation with human ratings of only 0.428. To address this issue, we propose VISJUDGE, a model specifically designed for visualization aesthetics and quality assessment. Experimental results demonstrate that VIS-JUDGE significantly narrows the gap with human judgment, reducing the MAE to 0.421 (a 23.9% reduction) and increasing the consistency with human experts to 0.687 (a 60.5% improvement) compared to GPT-5. The benchmark is available at https://github.com/HKUSTDial/VisJudgeBench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58ede094-d8f4-4dc3-b17e-81d0b102b791Cited by top-tier papers9
- DeepEye-SQL: A Software-Engineering-Inspired Text-to-SQL FrameworkBoyan Li, Chong Chen, Zhujun Xue, Yinan Mei et al.SIGMOD 2026 · 40 citations
- DAComp: Benchmarking Data Agents across the Full Data Intelligence LifecycleFangyu Lei, Jinxiang Meng, Yiming Huang, Junjie zhao et al.ICLR 2026 · 18 citations
- IGenBench: Benchmarking the Reliability of Text-to-Infographic GenerationYinghao Tang, Xueding Liu, Boyuan Zhang, Tingfeng Lan et al.ACL 2026 · 9 citations
- Document-to-Database: Extraction Meets Relational SemanticsZhengxuan Zhang, Zhuowen Liang, Jiazhuo Chen, Haixun Wang et al.VLDB 2026 · 3 citations
- Debugging Defective Visualizations: Empirical Insights Informing a Human-AI Co‑Debugging SystemShuyu Shen, Sirong Lu, Leixian Shen, Yuyu LuoCHI 2026 · 2 citations
Builds on21
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang et al.ICML 2024 · 345 citations
- Dashboard Design PatternsBenjamin Bach, Euan Freeman, Alfie Abdul-Rahman, Cagatay Turkay et al.IEEE VIS 2022 · 148 citations
- Natural Language to Visualization by Neural Machine TranslationYuyu Luo, Nan Tang, Guoliang Li, Jiawei Tang et al.IEEE VIS 2021 · 145 citations
- Composition and Configuration Patterns in Multiple-View VisualizationsXi Chen, Wei Zeng, Yanna Lin, Hayder Mahdi Al-Maneea et al.IEEE VIS 2020 · 138 citations
Related papers
- VisEval: A Benchmark for Data Visualization in the Era of Large Language ModelsNan Chen, Yuge Zhang, Jiahang Xu, Kan Ren et al.IEEE VIS 2024 · 44 citations
- Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from TextMizanur Rahman, Md. Tahmid Rahman Laskar, Shafiq Joty, Enamul HoqueEMNLP 2025 · 1 citation
- Can Vision-Language Models Assess Graphic Design Aesthetics? A Benchmark, Evaluation, and Dataset PerspectiveRuichuan An, Shizhao Sun, Danqing Huang, Mingxi Cheng et al.ICLR 2026 · 6 citations
- An Empirical Evaluation of the GPT-4 Multimodal Language Model on Visualization Literacy TasksAlexander Bendeck, John T. StaskoIEEE VIS 2024 · 40 citations
- Closing the Feedback Loop in Text2Vis: Refining Visualization with Vision-Language ModelsShengze Shi, Tao Ren, Guoliang Zhu, Guan Dong Feng et al.ACM MM 2025 · 2 citations
