Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL Benchmarks
Yuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai, Wenbo Li, Xuedi Qin
摘要
Natural language (nl) is a promising interaction paradigm for data visualization (vis). However, there are not any nl to vis (nl2vis) benchmarks available. Our goal is to provide the first nl2vis benchmark to enable and push the field of nl2vis, especially with deep learning technologies.
In this paper, we propose a nl2vis synthesizer (nl2sql-to-nl2vis) that synthesizes nl2vis benchmarks by piggybacking nl2sql benchmarks. The intuition is based on the semantic connection between sql queries and vis queries: sql queries specify what data is needed and vis queries additionally need to specify how to visualize. However, different from sql that has well-defined syntax, vis languages (e.g., Vega-Lite, VizQL, ggplot2) are syntactically very different. To provide nl2vis benchmarks that can support many vis languages, we use a unified intermediate representation, abstract syntax trees (ASTs), for both sql and vis queries. We can synthesize multiple vis trees through adding/deleting nodes to/from an sql tree. Each vis tree can then be converted to (any) vis language. The nl for vis will be modified based on the nl for sql to reflect corresponding tree edits.
We produce the first nl2vis benchmark (nvBench), by applying nl2sql-to-nl2vis on a popular nl2sql benchmark Spider, which covers 105 domains, supports seven common types of visualizations, and contains 25,750 (nl, vis) pairs. Our method reduces the manhour to 5.7% of developing a nl2vis benchmark from scratch (or building a nl2vis benchmark from scratch takes 17.5 × man-hours of our method). Extensive human validation, through 23 experts and 312 crowd workers, demonstrates the high-quality of nvBench.
In order to verify that nvBench can enable learning-based approaches, we develop a seq2vis model. Our experimental results show that seq2vis works well and significantly outperforms the state-of-the-art methods of the nl2vis task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- Natural Language to Visualization by Neural Machine TranslationYuyu Luo, Nan Tang, Guoliang Li, Jiawei Tang 等IEEE VIS 2021 · 被引用 145 次
- The Dawn of Natural Language to SQL: Are We Fully Ready? [Experiment, Analysis & Benchmark ]Boyan Li, Yuyu Luo, Chengliang Chai, Guoliang Li 等VLDB 2024 · 被引用 137 次
- Towards Natural Language-Based Visualization AuthoringYun Wang, Zhitao Hou, Leixian Shen, Tongshuang Wu 等IEEE VIS 2022 · 被引用 74 次
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li 等VLDB 2022 · 被引用 62 次
- HAIChart: Human and AI Paired Visualization SystemYupeng Xie, Yuyu Luo, Guoliang Li, Nan TangVLDB 2024 · 被引用 47 次
它引用的顶会 Paper8
- NL4DV: A Toolkit for Generating Analytic Specifications for Data Visualization from Natural Language QueriesArpit Narechania, Arjun Srinivasan, John T. StaskoIEEE VIS 2020 · 被引用 210 次
- Human-in-the-loop Outlier DetectionChengliang Chai, Lei Cao, Guoliang Li, Jian Li 等SIGMOD 2020 · 被引用 57 次
- IDEBench: A Benchmark for Interactive Data ExplorationPhilipp Eichmann, Emanuel Zgraggen, Carsten Binnig, Tim KraskaSIGMOD 2020 · 被引用 57 次
- RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL ParsersBailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov 等ACL 2020 · 被引用 39 次
- Interactive Cleaning for Progressive Visualization through Composite QuestionsYuyu Luo, Chengliang Chai, Xuedi Qin, Nan Tang 等ICDE 2020 · 被引用 37 次
相关 Paper
- RGVisNet: A Hybrid Retrieval-Generation Neural Framework Towards Automatic Data Visualization GenerationYuanfeng Song, Xuefang Zhao, Raymond Chi-Wing Wong, Di JiangKDD 2022 · 被引用 28 次
- Automated Data Visualization from Natural Language via Large Language Models: An Exploratory StudyYang Wu, Yao Wan, Hongyu Zhang, Yulei Sui 等SIGMOD 2024 · 被引用 44 次
- VisEval: A Benchmark for Data Visualization in the Era of Large Language ModelsNan Chen, Yuge Zhang, Jiahang Xu, Kan Ren 等IEEE VIS 2024 · 被引用 44 次
- Marrying Dialogue Systems with Data Visualization: Interactive Data Visualization Generation from Natural Language ConversationsYuanfeng Song, Xuefang Zhao, Raymond Chi-Wing WongKDD 2024 · 被引用 7 次
- nvAgent: Automated Data Visualization from Natural Language via Collaborative Agent WorkflowGeliang Ouyang, Jingyao Chen, Zhihe Nie, Yi Gui 等ACL 2025 · 被引用 22 次
