ACL2026
AnalystBench: Benchmarking professional long-form report generation with web-mined multimodal tasks
Chau Minh Pham, Zichao Wang, Puneet Mathur, Alexa Siu, Akriti Jain, Aparna Garimella, Ananya B. Sai, Nedim Lipka, Mohit Iyyer, Varun Manjunatha
Abstract
Large language models are increasingly used to draft long-form multimodal documents, but their end-to-end performance on professional report generation remains systematically understudied. We introduce AnalystBench, a continually extensible benchmark of 20 realworld report generation tasks grounded in multimodal document collections, where models must process millions of input tokens to produce long-form professional reports. Using expert-validated quality checklists and groundedness evaluation, we evaluate LLMs and coding agents and find that the best-performing model, GPT-5.1, scores highly on executive summary tasks (exceeding 90% on quality checklists) but degrades substantially on tasks requiring long-horizon synthesis over large inputs (down to 25-41%). Agent-based generation substantially benefits strong closed-source models like GPT-5.1, with checklist scores improving by 21.27 percentage points and visual coverage by 39 points over vanilla generation, but provides little benefit, and sometimes negative gains, for open-source models. like DeepSeek-R1 (-2.57 points). Expert reviewers note that while generated reports are grounded and clearly separate factual description from interpretation, they often fall short in actionability and precision, which highlights the remaining gap between system performance and professional needs. Task Domain Description S full R full Stext Rtext