SpreadsheetCoder: Formula Prediction from Semi-structured Context
Xinyun Chen, Petros Maniatis, Rishabh Singh, Charles Sutton, Hanjun Dai, Max Lin, Denny Zhou
摘要
Spreadsheet formula prediction has been an important program synthesis problem with many real-world applications. Previous works typically utilize input-output examples as the specification for spreadsheet formula synthesis, where each input-output pair simulates a separate row in the spreadsheet. However, this formulation does not fully capture the rich context in real-world spreadsheets. First, spreadsheet data entries are organized as tables, thus rows and columns are not necessarily independent from each other. In addition, many spreadsheet tables include headers, which provide high-level descriptions of the cell data. However, previous synthesis approaches do not consider headers as part of the specification. In this work, we present the first approach for synthesizing spreadsheet formulas from tabular context, which includes both headers and semi-structured tabular data. In particular, we propose SPREAD-SHEETCODER, a BERT-based model architecture to represent the tabular context in both row-based and column-based formats. We train our model on a large dataset of spreadsheets, and demonstrate that SPREADSHEETCODER achieves top-1 prediction accuracy of 42.51%, which is a considerable improvement over baselines that do not employ rich tabular context. Compared to the rule-based system, SPREADSHEETCODER assists 82% more users in composing formulas on Google Sheets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- LLMCompass: Enabling Efficient Hardware Design for Large Language Model InferenceHengrui Zhang, August Ning, Rohan Baskar Prabhakar, David WentzlaffISCA 2024 · 被引用 58 次
- FlashFill++: Scaling Programming by Example by Cutting to the ChaseJosé Cambronero, Sumit Gulwani, Vu Le, Daniel Perelman 等POPL 2023 · 被引用 27 次
- Natural Language to Code Generation in Interactive Data Science NotebooksPengcheng Yin, Wen-Ding Li, Kefan Xiao, Abhishek Rao 等ACL 2023 · 被引用 17 次
- SheetAgent: Towards a Generalist Agent for Spreadsheet Reasoning and Manipulation via Large Language ModelsYibin Chen, Yifu Yuan, Zeyu Zhang, Yan Zheng 等WWW 2025 · 被引用 13 次
- FormaT5: Abstention and Examples for Conditional Table Formatting with Natural LanguageMukul Singh, José Cambronero, Sumit Gulwani, Vu Le 等VLDB 2024 · 被引用 13 次
它引用的顶会 Paper3
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
- BUSTLE: Bottom-Up Program Synthesis Through Learning-Guided ExplorationAugustus Odena, Kensen Shi, David Bieber, Rishabh Singh 等ICLR 2021 · 被引用 60 次
- RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL ParsersBailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov 等ACL 2020 · 被引用 39 次
相关 Paper
- FORTAP: Using Formulas for Numerical-Reasoning-Aware Table PretrainingZhoujun Cheng, Haoyu Dong, Ran Jia, Pengfei Wu 等ACL 2022
- Semantic table structure identification in spreadsheetsYakun Zhang, Xiao Lv, Haoyu Dong, Wensheng Dou 等ISSTA 2021 · 被引用 11 次
- SheetPT: Spreadsheet Pre-training Based on Hierarchical Attention NetworkRan Jia, Qiyu Li, Zihan Xu, Xiaoyuan Jin 等AAAI 2023 · 被引用 3 次
- CORNET: Learning Table Formatting Rules By ExampleMukul Singh, José Pablo Cambronero Sánchez, Sumit Gulwani, Vu Le 等VLDB 2023 · 被引用 11 次
- HermEs: Interactive Spreadsheet Formula Prediction via Hierarchical Formulet ExpansionWanrong He, Haoyu Dong, Yihuai Gao, Zhichao Fan 等ACL 2023 · 被引用 5 次
