METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
Bingxuan Li, Yiwei Wang, Jiuxiang Gu, Kai-Wei Chang, Nanyun Peng
摘要
Chart generation aims to generate code to produce charts satisfying the desired visual properties, e.g., texts, layout, color, and type. It has great potential to empower the automatic professional report generation in financial analysis, research presentation, education, and healthcare. In this work, we build a vision-language model (VLM) based multi-agent framework for effective automatic chart generation. Generating high-quality charts requires both strong visual design skills and precise coding capabilities that embed the desired visual properties into code. Such a complex multi-modal reasoning process is difficult for direct prompting of VLMs. To resolve these challenges, we propose METAL (Multi-agEnT frAmework with vision Language models for chart generation), a multi-agent framework that decomposes the task of chart generation into the iterative collaboration among specialized agents. METAL achieves 5.2% improvement in accuracy over the current best result in the chart generation task. The METAL framework exhibits the phenomenon of test-time scaling: its performance increases monotonically as the logarithmic computational budget grows from 512 to 8192 tokens. In addition, we find that separating different modalities during the critique process of METAL boosts the self-correction capability of VLMs in the multimodal context.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time ScalingXinlei Yu, Chengming Xu, Zhangquan Chen, Yudong Zhang 等CVPR 2026 · 被引用 31 次
- Dual Latent Memory for Visual Multi-agent SystemXinlei Yu, Chengming Xu, Zhangquan Chen, Bo Yin 等ICML 2026 · 被引用 5 次
- EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound GenerationBingxuan Li, Yiming Cui, Yicheng He, Yiwei Wang 等CVPR 2026 · 被引用 5 次
- PEARL: Self-Evolving Assistant for Time Management with Reinforcement LearningBingxuan Li, Jeonghwan Kim, Cheng Qian, Xiusi Chen 等ACL 2026 · 被引用 2 次
- A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item GenerationSeonjeong Hwang, Jun Seo, Hyounghun Kim, Gary LeeACL 2026
它引用的顶会 Paper3
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code GenerationCheng Yang, Chufan Shi, Yaxin Liu, Bo Shui 等ICLR 2025 · 被引用 3 次
- Agents' Room: Narrative Generation through Multi-step CollaborationFantine Huot, Reinald Kim Amplayo, Jennimaria Palomaki, Alice Shoshana Jakobovits 等ICLR 2025
相关 Paper
- AMACE: Automatic Multi-Agent Chart Evolution for Iteratively Tailored Chart GenerationHyuk Namgoong, Jeesu Jung, Hyeonseok Kang, Yohan Lee 等EMNLP 2025
- FinSight: Towards Real-World Financial Deep ResearchJiajie Jin, Yuyao Zhang, Yimeng Xu, Hongjin Qian 等ACL 2026 · 被引用 5 次
- RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task EvaluationJiajun Zhang, Yuying Li, Zhixun Li, Xingyu Guo 等ACL 2026
- OneChart: Purify the Chart Structural Extraction via One Auxiliary TokenJinyue Chen, Lingyu Kong, Haoran Wei, Chenglong Liu 等ACM MM 2024 · 被引用 13 次
- Boosting Chart-to-Code Generation in MLLM via Dual Preference-Guided RefinementZhihan Zhang, Yixin Cao, Lizi LiaoACM MM 2025
