SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
Srivatsa Kundurthy, Clara Na, Michael Handley, Zach Kirshner, Chen Bo Calvin Zhang, Manasi Sharma, Emma Strubell, John Ling
摘要
We consider the task of end-to-end spreadsheet generation , where language models are prompted to produce spreadsheet artifacts to satisfy users' explicit and implicit constraints, specified in natural language. We introduce SpreadsheetArena , a platform for evaluating models' performance on the task via blind pairwise evaluations of LLM-generated spreadsheet workbooks. As with other complex, open-ended tasks, relevant evaluation criteria can vary greatly across use cases and prompts, often in ways that are difficult to formalize. Compared to general dialogue or text generation settings, spreadsheet generation presents unique challenges and opportunities: the task output structure is well-defined and multi-dimensional, and there are often complex interactivity and layout considerations. We observe that stylistic, structural, and functional features of preferred spreadsheets vary meaningfully across prompts. Expert evaluations of spreadsheets for finance prompts suggest that even highly ranked models do not reliably produce spreadsheets aligned with domain-specific best practices. We host a live arena and release a dataset of prompts, generated spreadsheets, and preference votes, which we hope will facilitate further study of tasks operating over spreadsheets as a challenging and interesting class of complex, open-ended tasks for LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- SheetCopilot: Bringing Software Productivity to the Next Level through Large Language ModelsHongxin Li, Jingran Su, Yuntao Chen, Qing Li 等NeurIPS 2023 · 被引用 75 次
相关 Paper
- GraphArena: Evaluating and Exploring Large Language Models on Graph ComputationJianheng Tang, Qifan Zhang, Yuhan Li, Nuo Chen 等ICLR 2025
- Copilot Arena: A Platform for Code LLM Evaluation in the WildWayne Chi, Valerie Chen, Anastasios Nikolas Angelopoulos, Wei-Lin Chiang 等ICML 2025
- A Benchmark and Framework for Evaluating Next Action Predictions in SpreadsheetsTejas Agrawal, Vu Le, Sumit Gulwani, Gust VerbruggenICML 2026
- LitReview Arena: Evaluating Literature Review Agents with Battle-style Peer Review PlatformRuotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue 等ICML 2026
- From Crowdsourced Data to High-quality Benchmarks: Arena-Hard and Benchbuilder PipelineTianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap 等ICML 2025
