FLAME: A Small Language Model for Spreadsheet Formulas
Harshit Joshi, Abishai Ebenezer, José Pablo Cambronero Sánchez, Sumit Gulwani, Aditya Kanade, Vu Le, Ivan Radicek, Gust Verbruggen
摘要
Spreadsheets are a vital tool for end-user data management. Using large language models for formula authoring assistance in these environments can be difficult, as these models are expensive to train and challenging to deploy due to their size (up to billions of parameters). We present FLAME, a transformer-based model trained exclusively on Excel formulas that leverages domain insights to achieve competitive performance while being substantially smaller (60M parameters) and training on two orders of magnitude less data. We curate a training dataset using sketch deduplication, introduce an Excel-specific formula tokenizer, and use domain-specific versions of masked span prediction and noisy auto-encoding as pre-training objectives. We evaluate FLAME on formula repair, formula completion, and similarity-based formula retrieval. FLAME can outperform much larger models, such as the Davinci (175B) and Cushman (12B) variants of Codex and CodeT5 (220M), in 10 of 14 evaluation settings for the repair and completion tasks. For formula retrieval, FLAME outperforms CodeT5, CodeBERT, and GraphCodeBERT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SheetCopilot: Bringing Software Productivity to the Next Level through Large Language ModelsHongxin Li, Jingran Su, Yuntao Chen, Qing Li 等NeurIPS 2023 · 被引用 75 次
- Encoding Spreadsheets for Large Language ModelsHaoyu Dong, Jianbo Zhao, Yuzhang Tian, Junyu Xiong 等EMNLP 2024 · 被引用 4 次
- Can an LLM Find Its Way Around a Spreadsheet?Cho-Ting Lee, Andrew Neeser, Shengzhe Xu, Jay Katyan 等ICSE 2025 · 被引用 1 次
- Learning from Near-Misses: Error-Aware Contrastive Few-Shot Learning for NL2FormulaZhihao Shuai, Yiyun Chen, Maolin Ma, Yutong Chen 等ACL 2026
- Unlocking SLM Potential for Data Analysis Code Generation via Non-Parametric Knowledge DistillationJinyang Li, Jack Williams, Nick McKenna, Arian Askari 等NeurIPS 2025
它引用的顶会 Paper20
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
相关 Paper
- From Rows to Reasoning: A Retrieval-Augmented Multimodal Framework for Spreadsheet UnderstandingAnmol Gulati, Sahil Sen, Waqar Sarguroh, Kevin PaulKDD 2026 · 被引用 6 次
- SpreadsheetCoder: Formula Prediction from Semi-structured ContextXinyun Chen, Petros Maniatis, Rishabh Singh, Charles Sutton 等ICML 2021 · 被引用 63 次
- Neurosymbolic repair for low-code formula languagesRohan Bavishi, Harshit Joshi, José Cambronero, Anna Fariha 等OOPSLA 2022 · 被引用 11 次
- PyDex: Repairing Bugs in Introductory Python Assignments using LLMsJialu Zhang, José Pablo Cambronero, Sumit Gulwani, Vu Le 等OOPSLA 2024 · 被引用 38 次
- FormaT5: Abstention and Examples for Conditional Table Formatting with Natural LanguageMukul Singh, José Cambronero, Sumit Gulwani, Vu Le 等VLDB 2024 · 被引用 13 次
