A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets
Tejas Agrawal, Vu Le, Sumit Gulwani, Gust Verbruggen
Abstract
Predictive code completion greatly accelerates how quickly developers work. In spreadsheets, despite being much more common, such auto-completion features are virtually non-existent. To address this gap, we introduce a benchmark for systems that observe a sequence of user actions in a spreadsheet and predict future actions. Two challenges are (1) the absence of edit histories in public spreadsheet corpora and (2) the complex space of spreadsheet actions (spatial, temporal, composite). To address (1), we manually curate 52 sequences of 12K actions that recreate spreadsheets from public corpora, seeded by parametrized heuristics and LLM refinement. To address (2), we propose an online evaluation that expects a prediction after each user action, accepts or rejects that prediction, updates the future actions upon acceptance, and repeats this until the target spreadsheet is obtained. We use multiple baseline predictors (including zero-shot LLMs, fine-tuned SLMs, and classical models) and analyze different properties that our benchmark teaches us, including but not limited to: properties of saved actions and false positives, efficiency, effect of user profiles, effect of triggers, and effect of context.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 33c690f5-7c0d-435a-ad93-56af0dc509cdBuilds on6
- On the Robustness of Code Generation Techniques: An Empirical Study on GitHub CopilotAntonio Mastropaolo, Luca Pascarella, Emanuela Guglielmi, Matteo Ciniselli et al.ICSE 2023 · 124 citations
- SheetCopilot: Bringing Software Productivity to the Next Level through Large Language ModelsHongxin Li, Jingran Su, Yuntao Chen, Qing Li et al.NeurIPS 2023 · 75 citations
- Language Models for Code Completion: A Practical EvaluationMaliheh Izadi, Jonathan Katzy, Tim van Dam, Marc Otten et al.ICSE 2024 · 51 citations
- When to Show a Suggestion? Integrating Human Feedback in AI-Assisted ProgrammingHussein Mozannar, Gagan Bansal, Adam Fourney, Eric HorvitzAAAI 2024 · 39 citations
- FlashFill++: Scaling Programming by Example by Cutting to the ChaseJosé Cambronero, Sumit Gulwani, Vu Le, Daniel Perelman et al.POPL 2023 · 27 citations
Related papers
- SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet WorkbooksSrivatsa Kundurthy, Clara Na, Michael Handley, Zach Kirshner et al.ICML 2026 · 2 citations
- ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex CodeJia Feng, Jiachen Liu, Cuiyun Gao, Chun Yong Chong et al.ASE 2024 · 7 citations
- EditBench: Evaluating LLM Abilities to Perform Real-World Instructed Code EditsWayne Chi, Valerie Chen, Ryan Shar, Aditya Mittal et al.ICLR 2026 · 7 citations
- M2RC-EVAL: Massively Multilingual Repository-level Code Completion EvaluationJiaheng Liu, Ken Deng, Congnan Liu, Jian Yang et al.ACL 2025 · 19 citations
- RepoBench: Benchmarking Repository-Level Code Auto-Completion SystemsTianyang Liu, Canwen Xu, Julian J. McAuleyICLR 2024 · 338 citations
