RGVisNet: A Hybrid Retrieval-Generation Neural Framework Towards Automatic Data Visualization Generation
Yuanfeng Song, Xuefang Zhao, Raymond Chi-Wing Wong, Di Jiang
Abstract
Recent years have witnessed the burgeoning of data visualization (DV) systems in both the research and the industrial communities since they provide vivid and powerful tools to convey the insights behind the massive data. A necessary step to visualize data is through creating suitable specifications in some declarative visualization languages (DVLs, e.g., Vega-Lite, ECharts). Due to the steep learning curve of mastering DVLs, automatically generating DVs via natural language questions, or text-to-vis, has been proposed and received great attention. However, existing neural network-based text-to-vis models, such as Seq2Vis or ncNet, usually generate DVs from scratch, limiting their performance due to the complex nature of this problem. Inspired by how developers reuse previously validated source code snippets from code search engines or a large-scale codebase when they conduct software development, we provide a novel hybrid retrieval-generation framework named RGVisNet for text-to-vis. It retrieves the most relevant DV query candidate as a prototype from the DV query codebase, and then revises the prototype to generate the desired DV query. Specifically, the DV query retrieval model is a neural ranking model which employs a schema-aware encoder for the NL question, and a GNN-based DV query encoder to capture the structure information of a DV query. At the same time, the DV query revision model shares the same structure and parameters of the encoders, and employs a DV grammar-aware decoder to reuse the retrieved prototype. Experimental evaluation on the public NVBench dataset validates that RGVisNet can significantly outperform existing generative text-to-vis models such as ncNet, by up to 74.28% relative improvement in terms of overall accuracy. To the best of our knowledge, RGVisNet is the first framework that seamlessly integrates the retrieval- with the generative-based approach for the text-to-vis task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Automated Data Visualization from Natural Language via Large Language Models: An Exploratory StudyYang Wu, Yao Wan, Hongyu Zhang, Yulei Sui et al.SIGMOD 2024 · 44 citations
- VisEval: A Benchmark for Data Visualization in the Era of Large Language ModelsNan Chen, Yuge Zhang, Jiahang Xu, Kan Ren et al.IEEE VIS 2024 · 44 citations
- nvAgent: Automated Data Visualization from Natural Language via Collaborative Agent WorkflowGeliang Ouyang, Jingyao Chen, Zhihe Nie, Yi Gui et al.ACL 2025 · 22 citations
- MultiVis-Agent: A Multi-Agent Framework with Logic Rules for Reliable and Comprehensive Cross-Modal Data VisualizationJinwei Lu, Yuanfeng Song, Chen Zhang, Raymond Chi-Wing WongSIGMOD 2026 · 14 citations
- Marrying Dialogue Systems with Data Visualization: Interactive Data Visualization Generation from Natural Language ConversationsYuanfeng Song, Xuefang Zhao, Raymond Chi-Wing WongKDD 2024 · 7 citations
Builds on9
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 1,586 citations
- Handling Information Loss of Graph Neural Networks for Session-based RecommendationTianwen Chen, Raymond Chi-Wing WongKDD 2020 · 292 citations
- Retrieval-based neural source code summarizationJian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun et al.ICSE 2020 · 242 citations
- NL4DV: A Toolkit for Generating Analytic Specifications for Data Visualization from Natural Language QueriesArpit Narechania, Arjun Srinivasan, John T. StaskoIEEE VIS 2020 · 210 citations
- Natural Language to Visualization by Neural Machine TranslationYuyu Luo, Nan Tang, Guoliang Li, Jiawei Tang et al.IEEE VIS 2021 · 145 citations
Related papers
- Towards Robustness of Text-to-Visualization Translation Against Lexical and Phrasal VariabilityJinwei Lu, Yuanfeng Song, Haodi Zhang, Chen Jason Zhang et al.ICDE 2025 · 3 citations
- Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL BenchmarksYuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai et al.SIGMOD 2021 · 90 citations
- FeVisQA: Free-Form Question Answering over Data VisualizationsYuanfeng Song, Jinwei Lu, Yuanwei Song, Caleb Chen Cao et al.ICDE 2025 · 2 citations
- Closing the Feedback Loop in Text2Vis: Refining Visualization with Vision-Language ModelsShengze Shi, Tao Ren, Guoliang Zhu, Guan Dong Feng et al.ACM MM 2025 · 2 citations
- SD-DVQ: Semantic-Driven Data Visualization Query Generation from Speech QueriesHaodi Zhang, Xiaohui Tang, Yuanfeng SongKDD 2026
