PlotCraft: Pushing the Limits of LLMs for Complex and Interactive Data Visualization
Jiajun Zhang, Jianke Zhang, Zeyu Cui, Jiaxi Yang, Lei Zhang, Zilei Wang, Qiang Liu, Liang Wang, Binyuan Hui, Junyang Lin
Abstract
Recent Large Language Models (LLMs) have demonstrated remarkable proficiency in code generation. However, their ability to create complex visualizations for scaled and structured data remains largely unevaluated and underdeveloped. To address this gap, we introduce PlotCraft , a new benchmark featuring 1k challenging visualization tasks that cover a wide range of topics, such as finance, scientific research, and sociology. The benchmark is structured around seven high-level visualization tasks and encompasses 48 distinct chart types. Crucially, it is the first to systematically evaluate both single-turn generation and multi-turn refinement across a diverse spectrum of task complexities. Our comprehensive evaluation of 23 leading LLMs on PlotCraft reveals obvious performance deficiencies in handling sophisticated visualization tasks. To bridge this performance gap, we develope SynthVis-30K , a large-scale, high-quality dataset of complex visualization code synthesized via a collaborative agent framework. Building upon this dataset, we develope PlotCraftor , a novel code generation model that achieves strong capabilities in complex data visualization with a remarkably small size. Across VisEval, PandasPlotBench, and our proposed PlotCraft, PlotCraftor shows performance comparable to that of leading proprietary approaches. Especially, on hard task, Our model achieves over 50% performance improvement. We will release the benchmark, dataset, and code at PlotCraft anonymous repository.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3793eb1b-2dda-448f-bfc8-32738eba09baCited by top-tier papers5
- Photon: Speedup Volume Understanding with Efficient Multimodal Large Language ModelsChengyu Fang, Heng Guo, Zheng Jiang, Chunming He et al.ICLR 2026 · 10 citations
- What Should I Cite? A RAG Benchmark for Academic Citation PredictionLeqi Zheng, Jiajun Zhang, Canzhi Chen, Chaokun Wang et al.WWW 2026 · 2 citations
- Scaling Agentic Verifier for Competitive CodingZeyao Ma, Jing Zhang, Xiaokang Zhang, Jiaxi Yang et al.ICML 2026 · 2 citations
- DV-World: Benchmarking Data Visualization Agents in Real-World ScenariosJinxiang Meng, Shaoping Huang, Fangyu Lei, Jingyu Guo et al.ICML 2026 · 1 citation
- RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task EvaluationJiajun Zhang, Yuying Li, Zhixun Li, Xingyu Guo et al.ACL 2026
Builds on8
- ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code GenerationXuanle Zhao, Xianzhen Luo, Qi Shi, Chi Chen et al.ACL 2025 · 62 citations
- VisEval: A Benchmark for Data Visualization in the Era of Large Language ModelsNan Chen, Yuge Zhang, Jiahang Xu, Kan Ren et al.IEEE VIS 2024 · 44 citations
- Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle TasksLinyuan Gong, Sida Wang, Mostafa Elhoushi, Alvin CheungICML 2024 · 33 citations
- Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction TuningXingchen Zeng, Haichuan Lin, Yilin Ye, Wei ZengIEEE VIS 2024 · 23 citations
- ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code GenerationCheng Yang, Chufan Shi, Yaxin Liu, Bo Shui et al.ICLR 2025 · 3 citations
Related papers
- Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive DashboardsTianhao Niu, Ziyu Han, Qiguang Chen, Shiqi Zhou et al.ACL 2026
- Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from TextMizanur Rahman, Md. Tahmid Rahman Laskar, Shafiq Joty, Enamul HoqueEMNLP 2025 · 1 citation
- VisCoder2: Building Multi-Language Visualization Coding AgentsYuansheng Ni, Songcheng Cai, Xiangchao Chen, Jiarong Liang et al.ICLR 2026 · 3 citations
- PlotCoder: Hierarchical Decoding for Synthesizing Visualization Code in Programmatic ContextXinyun Chen, Linyuan Gong, Alvin Cheung, Dawn SongACL 2021
- Text2Chart31: Instruction Tuning for Chart Generation with Automatic FeedbackFatemeh Pesaran Zadeh, Juyeon Kim, Jin-Hwa Kim, Gunhee KimEMNLP 2024 · 2 citations
