Chart2Code53: A Large-Scale Diverse and Complex Dataset for Enhancing Chart-to-Code Generation
Tianhao Niu, Yiming Cui, Baoxin Wang, Xiao Xu, Xin Yao, Qingfu Zhu, Dayong Wu, Shijin Wang, Wanxiang Che
摘要
Chart2Code has recently received significant attention in the multimodal community due to its potential to reduce the burden of visualization and promote a more detailed understanding of charts. However, existing Chart2Coderelated training datasets suffer from at least one of the following issues: (1) limited scale, (2) limited type coverage, and (3) inadequate complexity. To address these challenges, we seek more diverse sources that better align with real-world user distributions and propose dual data synthesis pipelines: (1) Synthesize based on online plotting code. ( 2 ) Synthesize based on the chart images in the academic paper. We create a large-scale Chart2Code training dataset Chart2Code53, including 53 chart types, 130K Chart-code pairs based on the pipeline. Experimental results demonstrate that even with few parameters, the model finetuned on Chart2Code53 achieves state-ofthe-art performance on multiple Chart2Code benchmarks within open-source models 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and EditingXiaoyan Su, Peijie Dong, Zhenheng Tang, Song Tang 等ICML 2026 · 被引用 3 次
- CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven EvolutionXiangxi Zheng, Kuang He, Jiayi Hu, Ping Yu 等ACL 2026 · 被引用 1 次
- Aligned Multi-View Scripts for Universal Chart-to-Code GenerationZhihan Zhang, Lizi LiaoACL 2026 · 被引用 1 次
它引用的顶会 Paper10
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code GenerationXuanle Zhao, Xianzhen Luo, Qi Shi, Chi Chen 等ACL 2025 · 被引用 62 次
- Visual Classification via Description from Large Language ModelsSachit Menon, Carl VondrickICLR 2023 · 被引用 57 次
- Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language ModelsLei Li, Yuqi Wang, Runxin Xu, Peiyi Wang 等ACL 2024 · 被引用 16 次
- TinyChart: Efficient Chart Understanding with Program-of-Thoughts Learning and Visual Token MergingLiang Zhang, Anwen Hu, Haiyang Xu, Ming Yan 等EMNLP 2024 · 被引用 15 次
相关 Paper
- ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart UnderstandingJovana Kondic, Pengyuan Li, Dhiraj Joshi, Isaac Sanchez 等CVPR 2026 · 被引用 7 次
- RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task EvaluationJiajun Zhang, Yuying Li, Zhixun Li, Xingyu Guo 等ACL 2026
- From Charts to Code: A Hierarchical Benchmark for Multimodal ModelsJiahao Tang, Henry Hengyuan Zhao, Lijian Wu, Zijian Zhang 等ACL 2026 · 被引用 5 次
- Effective Training Data Synthesis for Improving MLLM Chart UnderstandingYuwei Yang, Zeyu Zhang, Yunzhong Hou, Zhuowan Li 等ICCV 2025 · 被引用 4 次
- Text2Chart31: Instruction Tuning for Chart Generation with Automatic FeedbackFatemeh Pesaran Zadeh, Juyeon Kim, Jin-Hwa Kim, Gunhee KimEMNLP 2024 · 被引用 2 次
