Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL Benchmarks
Yuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai, Wenbo Li, Xuedi Qin
Abstract
Natural language (nl) is a promising interaction paradigm for data visualization (vis). However, there are not any nl to vis (nl2vis) benchmarks available. Our goal is to provide the first nl2vis benchmark to enable and push the field of nl2vis, especially with deep learning technologies.
In this paper, we propose a nl2vis synthesizer (nl2sql-to-nl2vis) that synthesizes nl2vis benchmarks by piggybacking nl2sql benchmarks. The intuition is based on the semantic connection between sql queries and vis queries: sql queries specify what data is needed and vis queries additionally need to specify how to visualize. However, different from sql that has well-defined syntax, vis languages (e.g., Vega-Lite, VizQL, ggplot2) are syntactically very different. To provide nl2vis benchmarks that can support many vis languages, we use a unified intermediate representation, abstract syntax trees (ASTs), for both sql and vis queries. We can synthesize multiple vis trees through adding/deleting nodes to/from an sql tree. Each vis tree can then be converted to (any) vis language. The nl for vis will be modified based on the nl for sql to reflect corresponding tree edits.
We produce the first nl2vis benchmark (nvBench), by applying nl2sql-to-nl2vis on a popular nl2sql benchmark Spider, which covers 105 domains, supports seven common types of visualizations, and contains 25,750 (nl, vis) pairs. Our method reduces the manhour to 5.7% of developing a nl2vis benchmark from scratch (or building a nl2vis benchmark from scratch takes 17.5 × man-hours of our method). Extensive human validation, through 23 experts and 312 crowd workers, demonstrates the high-quality of nvBench.
In order to verify that nvBench can enable learning-based approaches, we develop a seq2vis model. Our experimental results show that seq2vis works well and significantly outperforms the state-of-the-art methods of the nl2vis task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff8f8b37-deef-42ce-8008-73bcfb31a1b8Cited by top-tier papers35
- Natural Language to Visualization by Neural Machine TranslationYuyu Luo, Nan Tang, Guoliang Li, Jiawei Tang et al.IEEE VIS 2021 · 145 citations
- The Dawn of Natural Language to SQL: Are We Fully Ready? [Experiment, Analysis & Benchmark ]Boyan Li, Yuyu Luo, Chengliang Chai, Guoliang Li et al.VLDB 2024 · 137 citations
- Towards Natural Language-Based Visualization AuthoringYun Wang, Zhitao Hou, Leixian Shen, Tongshuang Wu et al.IEEE VIS 2022 · 74 citations
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li et al.VLDB 2022 · 62 citations
- HAIChart: Human and AI Paired Visualization SystemYupeng Xie, Yuyu Luo, Guoliang Li, Nan TangVLDB 2024 · 47 citations
Builds on8
- NL4DV: A Toolkit for Generating Analytic Specifications for Data Visualization from Natural Language QueriesArpit Narechania, Arjun Srinivasan, John T. StaskoIEEE VIS 2020 · 210 citations
- Human-in-the-loop Outlier DetectionChengliang Chai, Lei Cao, Guoliang Li, Jian Li et al.SIGMOD 2020 · 57 citations
- IDEBench: A Benchmark for Interactive Data ExplorationPhilipp Eichmann, Emanuel Zgraggen, Carsten Binnig, Tim KraskaSIGMOD 2020 · 57 citations
- RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL ParsersBailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov et al.ACL 2020 · 39 citations
- Interactive Cleaning for Progressive Visualization through Composite QuestionsYuyu Luo, Chengliang Chai, Xuedi Qin, Nan Tang et al.ICDE 2020 · 37 citations
Related papers
- RGVisNet: A Hybrid Retrieval-Generation Neural Framework Towards Automatic Data Visualization GenerationYuanfeng Song, Xuefang Zhao, Raymond Chi-Wing Wong, Di JiangKDD 2022 · 28 citations
- Automated Data Visualization from Natural Language via Large Language Models: An Exploratory StudyYang Wu, Yao Wan, Hongyu Zhang, Yulei Sui et al.SIGMOD 2024 · 44 citations
- VisEval: A Benchmark for Data Visualization in the Era of Large Language ModelsNan Chen, Yuge Zhang, Jiahang Xu, Kan Ren et al.IEEE VIS 2024 · 44 citations
- Marrying Dialogue Systems with Data Visualization: Interactive Data Visualization Generation from Natural Language ConversationsYuanfeng Song, Xuefang Zhao, Raymond Chi-Wing WongKDD 2024 · 7 citations
- nvAgent: Automated Data Visualization from Natural Language via Collaborative Agent WorkflowGeliang Ouyang, Jingyao Chen, Zhihe Nie, Yi Gui et al.ACL 2025 · 22 citations
