Lune

ACM MM2025Top-tier venue

Closing the Feedback Loop in Text2Vis: Refining Visualization with Vision-Language Models

Shengze Shi, Tao Ren, Guoliang Zhu, Guan Dong Feng, Jun Hu

2025Year
2Citations

Abstract

Text-to-Visualization (Text2Vis) generates data visualizations directly from natural language queries, democratizing access to data insights. Early Text2Vis efforts, primarily relying on rule-based systems and machine learning models, struggled to handle semantically intricate queries. The advent of large language models (LLMs) allows for better generalization in generating visualization code. However, LLM-based approaches have mainly focused on textual or code-level optimizations, neglecting the potential benefits of assessing and improving visualized charts. Hence, we propose Visualization Refinement (VisRef), a novel framework based on vision-language models (VLMs) to enhance Text2Vis outputs. (1) Knowledge Extraction -- VisRef extracts visualization assessment knowledge through a hierarchical contrastive prompt and multi-granularity quality assessment framework by comparing superior ground-truth charts with inferior Text2Vis outputs; and (2) VLM Fine-Tuning -- This knowledge is used to fine-tune a VLM through a two-stage approach, including warm-up and iterative preference alignment phases, to judge visualization quality and provide code-level refinement suggestions. Experimental results demonstrate that VisRef significantly outperforms state-of-the-art approaches, including LLM-based and VLM-prompted, and exhibits strong orthogonal compatibility with existing approaches.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines