ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
Muye Huang, Lingling Zhang, Jie Ma, Han Lai, Fangzhi Xu, Yifei Li, Wenjun Wu, Yaqiang Wu, Jun Liu
Abstract
Charts are high-density visualization carriers for complex data, serving as a crucial medium for information extraction and analysis. Automated chart understanding poses significant challenges to existing multimodal large language models (MLLMs) due to the need for precise and complex visual reasoning. Current step-by-step reasoning models primarily focus on text-based logical reasoning for chart understanding. However, they struggle to refine or correct their reasoning when errors stem from flawed visual understanding, as they lack the ability to leverage multimodal interaction for deeper comprehension. Inspired by human cognitive behavior, we propose ChartSketcher, a multimodal feedback-driven step-by-step reasoning method designed to address these limitations. ChartSketcher is a chart understanding model that employs Sketch-CoT, enabling MLLMs to annotate intermediate reasoning steps directly onto charts using a programmatic sketching library, iteratively feeding these visual annotations back into the reasoning process. This mechanism enables the model to visually ground its reasoning and refine its understanding over multiple steps. We employ a two-stage training strategy: a cold start phase to learn sketch-based reasoning patterns, followed by off-policy reinforcement learning to enhance reflection and generalization. Experiments demonstrate that ChartSketcher achieves promising performance on chart understanding benchmarks and general vision tasks, providing an interactive and interpretable approach to chart comprehension.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bbed2b8c-b265-4ef6-8bb8-8f93660a2dbcCited by top-tier papers6
- Rethinking Verification for LLM Code Generation: From Generation to TestingZihan Ma, Taolin Zhang, Maosong Cao, Junnan Liu et al.NeurIPS 2025 · 19 citations
- SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and MoreMuye Huang, Lingling Zhang, Yifei Li, Yaqiang Wu et al.CVPR 2026 · 7 citations
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal ModelsJingxuan Wei, Caijun Jia, Xi Bai, Xinglong Xu et al.CVPR 2026 · 7 citations
- See Less, See Right: Bi-directional Perceptual Shaping For Multimodal ReasoningShuoshuo Zhang, Yizhen Zhang, Jingjing Fu, Lei Song et al.CVPR 2026 · 3 citations
- Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense ChartsHongkun Pan, Yuwei Wu, Wanyi Hong, Shenghui Hu et al.CVPR 2026 · 1 citation
Builds on17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth et al.NeurIPS 2024 · 373 citations
Related papers
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart ReasoningZhengzhuo Xu, Sinan Du, Yiyan Qi, Siwen Lu et al.ICCV 2025 · 1 citation
- ChartR: Evaluating Reasoning Accuracy and Robustness in Chart Question AnsweringXiaojun Chen, Sixiao Luo, Ziqi Liu, Min Yang et al.CVPR 2026
- Boosting Chart-to-Code Generation in MLLM via Dual Preference-Guided RefinementZhihan Zhang, Yixin Cao, Lizi LiaoACM MM 2025
- Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training FrameworkYuchen He, Peizhi Ying, Liqi Cheng, Kuilin Peng et al.CHI 2026 · 1 citation
- Composition-Grounded Data Synthesis for Visual ReasoningXinyi Gu, Jiayuan Mao, Zhang-Wei Hong, Zhuoran Yu et al.ICLR 2026 · 1 citation
