Can an LLM Find Its Way Around a Spreadsheet?
Cho-Ting Lee, Andrew Neeser, Shengzhe Xu, Jay Katyan, Patrick Cross, Sharanya Pathakota, Marigold Norman, John Simeone, Jaganmohan Chandrasekaran, Naren Ramakrishnan
摘要
Spreadsheets are routinely used in business and scientific contexts, and one of the most vexing challenges data analysts face is performing data cleaning prior to analysis and evaluation. The ad-hoc and arbitrary nature of data cleaning problems, such as typos, inconsistent formatting, missing values, and a lack of standardization, often creates the need for highly specialized pipelines. We ask whether an LLM can find its way around a spreadsheet and how to support end-users in taking their free-form data processing requests to fruition. Just like RAG retrieves context to answer users' queries, we demonstrate how we can retrieve elements from a code library to compose data processing pipelines. Through comprehensive experiments, we demonstrate the quality of our system and how it is able to continuously augment its vocabulary by saving new codes and pipelines back to the code library for future retrieval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Jigsaw: Large Language Models meet Program SynthesisNaman Jain, Skanda Vaidyanath, Arun Iyer, Nagarajan Natarajan 等ICSE 2022 · 被引用 134 次
- SkCoder: A Sketch-based Approach for Automatic Code GenerationJia Li, Yongmin Li, Ge Li, Zhi Jin 等ICSE 2023 · 被引用 50 次
- FLAME: A Small Language Model for Spreadsheet FormulasHarshit Joshi, Abishai Ebenezer, José Pablo Cambronero Sánchez, Sumit Gulwani 等AAAI 2024 · 被引用 21 次
相关 Paper
- Reliable and Cost-Effective Exploratory Data Analysis via Graph-Guided RAGMossad Helali, Yutai Luo, Tae Jun Ham, Jim Plotts 等EMNLP 2025
- RAG Without the Lag: Enabling "What-If" Analysis for Retrieval-Augmented Generation PipelinesQuentin Romero Lauro, Shreya Shankar, Sepanta Zeighami, Aditya G. ParameswaranCHI 2026 · 被引用 1 次
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data LakesSimran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan 等VLDB 2024 · 被引用 165 次
- ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User PromptsJiajun Zhu, Xinyu Cheng, Zhongsu Luo, Yunfan Zhou 等UIST 2025 · 被引用 1 次
- Dango: A Mixed-Initiative Data Wrangling System using Large Language ModelWei-Hao Chen, Weixi Tong, Amanda Case, Tianyi ZhangCHI 2025 · 被引用 19 次
