SD-DVQ: Semantic-Driven Data Visualization Query Generation from Speech Queries
Haodi Zhang, Xiaohui Tang, Yuanfeng Song
Abstract
In the era of big data, effectively leveraging vast amounts of data is one of the major challenges across industries, especially in data visualization, which helps decision-makers intuitively understand data. However, many existing data visualization tools require users to have certain technical skills, posing a barrier for users without a technical background. With the widespread use of voice interfaces, particularly smart assistants, the task of converting voice queries into data visualizations, known as ''Speech-to-Vis,'' has emerged. Yet, existing methods still face issues such as speech recognition errors and insufficient model generalization. We propose SD-DVQ, a Semantic-Driven Data Visualization Query Generation from Speech Queries framework, to address these issues. Unlike previous methods that rely on ASR-based cascades or completely end-to-end systems, SD-DVQ generates semantically consistent natural language queries from speech with schema information, and then improves downstream DVQ generation through schema filtering and database-feedback-based self-correction. Experiments on the SpeechNVBench benchmark show that SD-DVQ achieves 45.02% exact-match accuracy, outperforming the strongest baseline by 5.07 percentage points.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- RGVisNet: A Hybrid Retrieval-Generation Neural Framework Towards Automatic Data Visualization GenerationYuanfeng Song, Xuefang Zhao, Raymond Chi-Wing Wong, Di JiangKDD 2022 · 28 citations
- Marrying Dialogue Systems with Data Visualization: Interactive Data Visualization Generation from Natural Language ConversationsYuanfeng Song, Xuefang Zhao, Raymond Chi-Wing WongKDD 2024 · 7 citations
- Towards Robustness of Text-to-Visualization Translation Against Lexical and Phrasal VariabilityJinwei Lu, Yuanfeng Song, Haodi Zhang, Chen Jason Zhang et al.ICDE 2025 · 3 citations
- Closing the Feedback Loop in Text2Vis: Refining Visualization with Vision-Language ModelsShengze Shi, Tao Ren, Guoliang Zhu, Guan Dong Feng et al.ACM MM 2025 · 2 citations
- Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL BenchmarksYuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai et al.SIGMOD 2021 · 90 citations
