Infogen: Generating Complex Statistical Infographics from Documents
Akash Ghosh, Aparna Garimella, Pritika Ramu, Sambaran Bandyopadhyay, Sriparna Saha
Abstract
Statistical infographics are powerful tools that simplify complex data into visually engaging and easy-to-understand formats. Despite advancements in AI, particularly with LLMs, existing efforts have been limited to generating simple charts, with no prior work addressing the creation of complex infographics from textheavy documents that demand a deep understanding of the content. We address this gap by introducing the task of generating statistical infographics composed of multiple subcharts (e.g., line, bar, pie) that are contextually accurate, insightful, and visually aligned. To achieve this, we define infographic metadata, that includes its title and textual insights, along with sub-chart-specific details such as their corresponding data, alignment, etc. We also present Infodat, the first benchmark dataset for text-to-infographic metadata generation, where each sample links a document to its metadata. We propose Infogen, a two-stage framework where fine-tuned LLMs first generate metadata, which is then converted into infographic code. Extensive evaluations on Infodat demonstrate that Infogen achieves state-of-the-art performance, outperforming both closed and opensource LLMs in text-to-statistical infographic generation. The sample datapoints from Infodat can be accessed through this link * Work done during internship at Adobe Research. 1 While the term infographic can refer to a wide range of visual illustrations, we restrict ourselves to only those visuals
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 968145b0-0c05-4dbf-8b95-f8ec31aba77eCited by top-tier papers2
- CLINIC : Evaluating Multilingual Trustworthiness in Language Models for HealthcareAkash Ghosh, Srivarshinee Sridhar, Raghav Kaushik Ravi, Muhsin Muhsin et al.ICML 2026 · 6 citations
- Beyond Static Artifacts: An Evolutionary Framework for Synthetic Claim GenerationYeqing Teng, Jiasheng Si, Shuxia Lin, Linhai Zhang et al.ACL 2026
Builds on10
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- ReFT: Representation Finetuning for Language ModelsZhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger et al.NeurIPS 2024 · 233 citations
- Calliope: Automatic Visual Data Story Generation from a SpreadsheetDanqing Shi, Xinyue Xu, Fuling Sun, Yang Shi et al.IEEE VIS 2020 · 179 citations
- UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and ReasoningAhmed Masry, Parsa Kavehzadeh, Do Xuan Long, Enamul Hoque et al.EMNLP 2023 · 48 citations
Related papers
- IGenBench: Benchmarking the Reliability of Text-to-Infographic GenerationYinghao Tang, Xueding Liu, Boyuan Zhang, Tingfeng Lan et al.ACL 2026 · 9 citations
- ChartGalaxy: A Dataset for Infographic Chart Understanding and GenerationZhen Li, Duan Li, Yukai Guo, Xinyuan Guo et al.ICLR 2026 · 16 citations
- InfoDet: A Dataset for Infographic Element DetectionJiangning Zhu, Yuxing Zhou, Zheng Wang, Juntao Yao et al.ICLR 2026 · 4 citations
- Doc2Chart: Intent-Driven Zero-Shot Chart Generation from DocumentsAkriti Jain, Pritika Ramu, Aparna Garimella, Apoorv SaxenaEMNLP 2025
- Retrieve-Then-Adapt: Example-based Automatic Generation for Proportion-related InfographicsChunyao Qian, Shizhao Sun, Weiwei Cui, Jian-Guang Lou et al.IEEE VIS 2020 · 53 citations
