SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
Srivatsa Kundurthy, Clara Na, Michael Handley, Zach Kirshner, Chen Bo Calvin Zhang, Manasi Sharma, Emma Strubell, John Ling
Abstract
We consider the task of end-to-end spreadsheet generation , where language models are prompted to produce spreadsheet artifacts to satisfy users' explicit and implicit constraints, specified in natural language. We introduce SpreadsheetArena , a platform for evaluating models' performance on the task via blind pairwise evaluations of LLM-generated spreadsheet workbooks. As with other complex, open-ended tasks, relevant evaluation criteria can vary greatly across use cases and prompts, often in ways that are difficult to formalize. Compared to general dialogue or text generation settings, spreadsheet generation presents unique challenges and opportunities: the task output structure is well-defined and multi-dimensional, and there are often complex interactivity and layout considerations. We observe that stylistic, structural, and functional features of preferred spreadsheets vary meaningfully across prompts. Expert evaluations of spreadsheets for finance prompts suggest that even highly ranked models do not reliably produce spreadsheets aligned with domain-specific best practices. We host a live arena and release a dataset of prompts, generated spreadsheets, and preference votes, which we hope will facilitate further study of tasks operating over spreadsheets as a challenging and interesting class of complex, open-ended tasks for LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ae96e33-9fcb-4c3e-8dd6-5342f222186dBuilds on13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- SheetCopilot: Bringing Software Productivity to the Next Level through Large Language ModelsHongxin Li, Jingran Su, Yuntao Chen, Qing Li et al.NeurIPS 2023 · 75 citations
Related papers
- GraphArena: Evaluating and Exploring Large Language Models on Graph ComputationJianheng Tang, Qifan Zhang, Yuhan Li, Nuo Chen et al.ICLR 2025
- Copilot Arena: A Platform for Code LLM Evaluation in the WildWayne Chi, Valerie Chen, Anastasios Nikolas Angelopoulos, Wei-Lin Chiang et al.ICML 2025
- A Benchmark and Framework for Evaluating Next Action Predictions in SpreadsheetsTejas Agrawal, Vu Le, Sumit Gulwani, Gust VerbruggenICML 2026
- LitReview Arena: Evaluating Literature Review Agents with Battle-style Peer Review PlatformRuotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue et al.ICML 2026
- From Crowdsourced Data to High-quality Benchmarks: Arena-Hard and Benchbuilder PipelineTianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap et al.ICML 2025
