Encoding Spreadsheets for Large Language Models
Haoyu Dong, Jianbo Zhao, Yuzhang Tian, Junyu Xiong, Mengyu Zhou, Yun Lin, José Cambronero, Yeye He, Shi Han, Dongmei Zhang
摘要
Spreadsheets are characterized by their extensive two-dimensional grids, flexible layouts, and varied formatting options, which pose significant challenges for large language models (LLMs). In response, we introduce SHEETENCODER, pioneering an efficient encoding method designed to unleash and optimize LLMs' powerful understanding and reasoning capability on spreadsheets. Initially, we propose a vanilla serialization approach that incorporates cell addresses, values, and formats. However, this approach was limited by LLMs' token constraints, making it impractical for most applications. To tackle this challenge, three innovative modules are proposed to compress spreadsheets effectively: structural-anchor-based compression, inverse index translation, and data-format-aware aggregation. It significantly improves performance in spreadsheet table detection task, outperforming the vanilla approach by 25.6% in GPT4's incontext learning setting. Moreover, fine-tuned LLM with SHEETENCODER has an average compression ratio of 25×, but achieves a stateof-the-art 78.9% F1 score, surpassing the best existing models by 12.3%, demonstrating that SHEETENCODER greatly boosts LLMs's performance on spreadsheet data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- TableMaster: A Recipe to Advance Table Understanding with Language ModelsLang Cao, Hanbing LiuICLR 2026 · 被引用 20 次
- ASTRA: Adaptive Semantic Tree Reasoning Architecture for Complex Table Question AnsweringXiaoke Guo, Songze Li, Zhiqiang Liu, Zhaoyan Gong 等ACL 2026 · 被引用 3 次
- Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMsEyal German, Sagiv Antebi, Daniel Samira, Asaf Shabtai 等ICLR 2026 · 被引用 3 次
- SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet WorkbooksSrivatsa Kundurthy, Clara Na, Michael Handley, Zach Kirshner 等ICML 2026 · 被引用 2 次
- TopBench: A Benchmark for Implicit Predictive Reasoning in Tabular Question AnsweringAn-Yang Ji, Jun-Peng Jiang, De-Chuan Zhan, Han-Jia YeICML 2026 · 被引用 1 次
它引用的顶会 Paper13
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi 等ICLR 2022 · 被引用 347 次
- Chain-of-Table: Evolving Tables in the Reasoning Chain for Table UnderstandingZilong Wang, Hao Zhang, Chun-Liang Li, Julian Martin Eisenschlos 等ICLR 2024 · 被引用 244 次
- StructGPT: A General Framework for Large Language Model to Reason over Structured DataJinhao Jiang, Kun Zhou, Zican Dong, Keming Ye 等EMNLP 2023 · 被引用 173 次
- Retrieval meets Long Context Large Language ModelsPeng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee 等ICLR 2024 · 被引用 131 次
相关 Paper
- SheetAgent: Towards a Generalist Agent for Spreadsheet Reasoning and Manipulation via Large Language ModelsYibin Chen, Yifu Yuan, Zeyu Zhang, Yan Zheng 等WWW 2025 · 被引用 13 次
- SheetBrain: A Neuro-Symbolic Agent for Accurate Reasoning over Complex and Large SpreadsheetsZiwei Wang, Jiayuan Su, Mengyu Zhou, Huaxing Zeng 等AAAI 2026 · 被引用 2 次
- Towards Robust Real-World Spreadsheet Understanding with Multi-Agent Multi-Format ReasoningHouxing Ren, Mingjie Zhan, Zimu Lu, Ke Wang 等ACL 2026
- SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based ReflectionQin Chen, Yuanyi Ren, Xiaojun Ma, Mugeng Liu 等EMNLP 2025
- SpreadsheetCoder: Formula Prediction from Semi-structured ContextXinyun Chen, Petros Maniatis, Rishabh Singh, Charles Sutton 等ICML 2021 · 被引用 63 次
