OneChart: Purify the Chart Structural Extraction via One Auxiliary Token
Jinyue Chen, Lingyu Kong, Haoran Wei, Chenglong Liu, Zheng Ge, Liang Zhao, Jianjian Sun, Chunrui Han, Xiangyu Zhang
Abstract
Chart parsing poses a significant challenge due to the diversity of styles, values, texts, and so forth. Even advanced large vision-language models (LVLMs) with billions of parameters struggle to handle such tasks satisfactorily. To address this, we propose OneChart: a reliable agent specifically devised for the structural extraction of chart information. Similar to popular LVLMs, OneChart incorporates an autoregressive main body. Uniquely, to enhance the reliability of the numerical parts of the output, we introduce an auxiliary token placed at the beginning of the total tokens along with an additional decoder. The numerically optimized (auxiliary) token allows subsequent tokens for chart parsing to capture enhanced numerical features through causal attention. Furthermore, with the aid of the auxiliary token, we have devised a self-evaluation mechanism that enables the model to gauge the reliability of its chart parsing results by providing confidence scores for the generated content. Compared to current state-of-the-art (SOTA) chart parsing models, e.g., DePlot, ChartVLM, ChartAst, OneChart significantly outperforms in Average Precision (AP) for chart structural extraction across multiple public benchmarks, despite enjoying only 0.2 billion parameters. Moreover, as a chart parsing agent, it also brings 10%+ accuracy gains for the popular LVLM (LLaVA-1.6) in the downstream ChartQA benchmark. Our code, model and datasets ara available at https://onechartt.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd8ad76d-2f0e-4f22-8862-1dbd179e355dCited by top-tier papers17
- CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in LiteracyZhibo Yang, Jun Tang, Zhaohai Li, Pengfei Wang et al.ICCV 2025 · 15 citations
- ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart UnderstandingMuye Huang, Lingling Zhang, Jie Ma, Han Lai et al.NeurIPS 2025 · 13 citations
- SmolDocling: An Ultra-Compact Vision-Language Model for End-To-End Multi-Modal Document ConversionAhmed S. Nassar, Matteo Omenetti, Maksym Lysak, Nikolaos Livathinos et al.ICCV 2025 · 9 citations
- VProChart: Answering Chart Question Through Visual Perception Alignment Agent and Programmatic Solution ReasoningMuye Huang, Lingling Zhang, Han Lai, Wenjun Wu et al.AAAI 2025 · 7 citations
- SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and MoreMuye Huang, Lingling Zhang, Yifei Li, Yaqiang Wu et al.CVPR 2026 · 7 citations
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- DreamLLM: Synergistic Multimodal Comprehension and CreationRunpei Dong, Chunrui Han, Yuang Peng, Zekun Qi et al.ICLR 2024 · 315 citations
Related papers
- ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question AnsweringJingxuan Wei, Nan Xu, Junnan Zhu, Yanni Hao et al.EMNLP 2025 · 6 citations
- Visual Self-Refine: A Pixel-Guided Paradigm for Accurate Chart ParsingJinsong Li, Xiaoyi Dong, Yuhang Zang, Yuhang Cao et al.ICLR 2026 · 6 citations
- ChartR: Evaluating Reasoning Accuracy and Robustness in Chart Question AnsweringXiaojun Chen, Sixiao Luo, Ziqi Liu, Min Yang et al.CVPR 2026
- METAL: A Multi-Agent Framework for Chart Generation with Test-Time ScalingBingxuan Li, Yiwei Wang, Jiuxiang Gu, Kai-Wei Chang et al.ACL 2025
- TinyChart: Efficient Chart Understanding with Program-of-Thoughts Learning and Visual Token MergingLiang Zhang, Anwen Hu, Haiyang Xu, Ming Yan et al.EMNLP 2024 · 15 citations
