A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets
Tejas Agrawal, Vu Le, Sumit Gulwani, Gust Verbruggen
摘要
Predictive code completion greatly accelerates how quickly developers work. In spreadsheets, despite being much more common, such auto-completion features are virtually non-existent. To address this gap, we introduce a benchmark for systems that observe a sequence of user actions in a spreadsheet and predict future actions. Two challenges are (1) the absence of edit histories in public spreadsheet corpora and (2) the complex space of spreadsheet actions (spatial, temporal, composite). To address (1), we manually curate 52 sequences of 12K actions that recreate spreadsheets from public corpora, seeded by parametrized heuristics and LLM refinement. To address (2), we propose an online evaluation that expects a prediction after each user action, accepts or rejects that prediction, updates the future actions upon acceptance, and repeats this until the target spreadsheet is obtained. We use multiple baseline predictors (including zero-shot LLMs, fine-tuned SLMs, and classical models) and analyze different properties that our benchmark teaches us, including but not limited to: properties of saved actions and false positives, efficiency, effect of user profiles, effect of triggers, and effect of context.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- On the Robustness of Code Generation Techniques: An Empirical Study on GitHub CopilotAntonio Mastropaolo, Luca Pascarella, Emanuela Guglielmi, Matteo Ciniselli 等ICSE 2023 · 被引用 124 次
- SheetCopilot: Bringing Software Productivity to the Next Level through Large Language ModelsHongxin Li, Jingran Su, Yuntao Chen, Qing Li 等NeurIPS 2023 · 被引用 75 次
- Language Models for Code Completion: A Practical EvaluationMaliheh Izadi, Jonathan Katzy, Tim van Dam, Marc Otten 等ICSE 2024 · 被引用 51 次
- When to Show a Suggestion? Integrating Human Feedback in AI-Assisted ProgrammingHussein Mozannar, Gagan Bansal, Adam Fourney, Eric HorvitzAAAI 2024 · 被引用 39 次
- FlashFill++: Scaling Programming by Example by Cutting to the ChaseJosé Cambronero, Sumit Gulwani, Vu Le, Daniel Perelman 等POPL 2023 · 被引用 27 次
相关 Paper
- SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet WorkbooksSrivatsa Kundurthy, Clara Na, Michael Handley, Zach Kirshner 等ICML 2026 · 被引用 2 次
- ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex CodeJia Feng, Jiachen Liu, Cuiyun Gao, Chun Yong Chong 等ASE 2024 · 被引用 7 次
- EditBench: Evaluating LLM Abilities to Perform Real-World Instructed Code EditsWayne Chi, Valerie Chen, Ryan Shar, Aditya Mittal 等ICLR 2026 · 被引用 7 次
- M2RC-EVAL: Massively Multilingual Repository-level Code Completion EvaluationJiaheng Liu, Ken Deng, Congnan Liu, Jian Yang 等ACL 2025 · 被引用 19 次
- RepoBench: Benchmarking Repository-Level Code Auto-Completion SystemsTianyang Liu, Canwen Xu, Julian J. McAuleyICLR 2024 · 被引用 338 次
