DataVisT5: A Pre-Trained Language Model for Jointly Understanding Text and Data Visualization
Zhuoyue Wan, Yuanfeng Song, Shuaimin Li, Chen Jason Zhang, Raymond Chi-Wing Wong
Abstract
Data visualization (DV) is the fundamental and premise tool to improve the efficiency in conveying the insights behind the big data, which has been widely accepted in existing data-driven world. Task automation in DV, such as converting natural language queries to visualizations (i.e., text-to-vis), gener-ating explanations from visualizations (i.e., vis-to-text), answering DV-related questions in free form (i.e. Fe VisQA), and explicating tabular data (i.e., table-to-text), is vital for advancing the field. Despite their potential, the application of pre-trained language models (PLMs) like T5 and BERT in DV has been limited by high costs and challenges in handling cross-modal information, leading to few studies on PLMs for DV. We introduce Data VisT5, a novel PLM tailored for DV that enhances the T5 architecture through a hybrid objective pre-training and multi-task fine-tuning strategy, integrating text and DV datasets to effectively interpret cross-modal semantics. Extensive evaluations on public datasets show that Data VisT5 consistently outperforms current state-of-the-art models and higher-parameter Large Language Models (LLMs) on various DV-related tasks. We anticipate that Data VisT5 will not only inspire further research on vertical PLMs but also expand the range of applications for PLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 703571ad-2004-430c-a2d3-85c1386436efCited by top-tier papers3
- nvAgent: Automated Data Visualization from Natural Language via Collaborative Agent WorkflowGeliang Ouyang, Jingyao Chen, Zhihe Nie, Yi Gui et al.ACL 2025 · 22 citations
- VisPoison: An Effective Backdoor Attack Framework for Tabular Data Visualization ModelsShuaimin Li, Chen Jason Zhang, Xuanang Chen, Anni Peng et al.ICDE 2026
- OsmT: Bridging Openstreetmap Queries and Natural Language With Open-Source Tag-Aware Language ModelsZhuoyue Wan, Wentao Hu, Chen Jason Zhang, Yuanfeng Song et al.ICDE 2026
Builds on18
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
Related papers
- Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from TextMizanur Rahman, Md. Tahmid Rahman Laskar, Shafiq Joty, Enamul HoqueEMNLP 2025 · 1 citation
- FeVisQA: Free-Form Question Answering over Data VisualizationsYuanfeng Song, Jinwei Lu, Yuanwei Song, Caleb Chen Cao et al.ICDE 2025 · 2 citations
- VisEval: A Benchmark for Data Visualization in the Era of Large Language ModelsNan Chen, Yuge Zhang, Jiahang Xu, Kan Ren et al.IEEE VIS 2024 · 44 citations
- Automated Data Visualization from Natural Language via Large Language Models: An Exploratory StudyYang Wu, Yao Wan, Hongyu Zhang, Yulei Sui et al.SIGMOD 2024 · 44 citations
- Marrying Dialogue Systems with Data Visualization: Interactive Data Visualization Generation from Natural Language ConversationsYuanfeng Song, Xuefang Zhao, Raymond Chi-Wing WongKDD 2024 · 7 citations
