ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering
Rachneet Kaur, Nishan Srishankar, Zhen Zeng, Sumitra Ganesh
摘要
Recent multimodal LLMs have shown promise in chart-based visual question answering, but their performance declines sharply on unannotated charts-those requiring precise visual interpretation rather than relying on textual shortcuts. To address this, we introduce ChartAgent, a novel agentic framework that explicitly performs visual reasoning directly within the chart's spatial domain. Unlike textual chain-of-thought reasoning, ChartAgent iteratively decomposes queries into visual subtasks and actively manipulates and interacts with chart images through specialized actions such as drawing annotations, cropping regions (e.g., segmenting pie slices, isolating bars), and localizing axes, using a library of chart-specific vision tools to fulfill each subtask. This iterative reasoning process closely mirrors human cognitive strategies for chart comprehension. ChartAgent achieves state-of-the-art accuracy on the ChartBench and ChartX benchmarks, surpassing prior methods by up to 16.07% absolute gain overall and 17.31% on unannotated, numerically intensive queries. Furthermore, our analyses show that ChartAgent is (a) effective across diverse chart types, (b) achieves the highest scores across varying visual and reasoning complexity levels, and (c) serves as a plug-and-play framework that boosts performance across diverse underlying LLMs. Our work is among the first to demonstrate visually grounded reasoning for chart understanding using tool-augmented multimodal agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Deep-Reporter: Deep Research for Grounded Multimodal Long-Form GenerationFangda Ye, Kuicai Dong, Zhifei Xie, Yuxin Hu 等ACL 2026 · 被引用 2 次
- Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart UnderstandingJianzhu Bao, Haozhen Zhang, Kuicai Dong, Bozhi Wu 等ACL 2026
它引用的顶会 Paper10
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong 等NeurIPS 2024 · 被引用 858 次
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 被引用 732 次
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun 等ICML 2024 · 被引用 496 次
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao 等ICLR 2026 · 被引用 321 次
相关 Paper
- ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question AnsweringJingxuan Wei, Nan Xu, Junnan Zhu, Yanni Hao 等EMNLP 2025 · 被引用 6 次
- VProChart: Answering Chart Question Through Visual Perception Alignment Agent and Programmatic Solution ReasoningMuye Huang, Lingling Zhang, Han Lai, Wenjun Wu 等AAAI 2025 · 被引用 7 次
- ChartR: Evaluating Reasoning Accuracy and Robustness in Chart Question AnsweringXiaojun Chen, Sixiao Luo, Ziqi Liu, Min Yang 等CVPR 2026
- DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense DocumentsZhuoran Yu, Le T Nguyen, Jaden Park, Xinyi Gu 等ICML 2026
- DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific ChartsYujing Lu, Ling Zhong, Jing Yang, Weiming Li 等AAAI 2026
