SheetPT: Spreadsheet Pre-training Based on Hierarchical Attention Network
Ran Jia, Qiyu Li, Zihan Xu, Xiaoyuan Jin, Lun Du, Haoyu Dong, Xiao Lv, Shi Han, Dongmei Zhang
Abstract
Spreadsheets are an important and unique type of business document for data storage, analysis and presentation. The distinction between spreadsheets and most other types of digital documents lies in that spreadsheets provide users with high flexibility of data organization on the grid. Existing related techniques mainly focus on the tabular data and are incompetent in understanding the entire sheet. On the one hand, spreadsheets have no explicit separation across tabular data and other information, leaving a gap for the deployment of such techniques. On the other hand, pervasive data dependence and semantic relations across the sheet require comprehensive modeling of all the information rather than only the tables. In this paper, we propose SheetPT, the first pre-training technique on spreadsheets to enable effective representation learning under this scenario. For computational effectiveness and efficiency, we propose the coherent chunk, an intermediate semantic unit of sheet structure; and we accordingly devise a hierarchical attention-based architecture to capture contextual information across different structural granularities. Three pre-training objectives are also designed to ensure sufficient training against millions of spreadsheets. Two representative downstream tasks, formula prediction and sheet structure recognition are utilized to evaluate its capability and the prominent results reveal its superiority over existing state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi et al.ICLR 2022 · 347 citations
- TUTA: Tree-based Transformers for Generally Structured Table Pre-trainingZhiruo Wang, Haoyu Dong, Ran Jia, Jia Li et al.KDD 2021 · 88 citations
- SpreadsheetCoder: Formula Prediction from Semi-structured ContextXinyun Chen, Petros Maniatis, Rishabh Singh, Charles Sutton et al.ICML 2021 · 63 citations
- TabularNet: A Neural Network Architecture for Understanding Semantic Structures of Tabular DataLun Du, Fei Gao, Xu Chen, Ran Jia et al.KDD 2021 · 58 citations
Related papers
- GetPt: Graph-enhanced General Table Pre-training with Alternate Attention NetworkRan Jia, Haoming Guo, Xiaoyuan Jin, Chao Yan et al.KDD 2023 · 3 citations
- FORTAP: Using Formulas for Numerical-Reasoning-Aware Table PretrainingZhoujun Cheng, Haoyu Dong, Ran Jia, Pengfei Wu et al.ACL 2022
- SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based ReflectionQin Chen, Yuanyi Ren, Xiaojun Ma, Mugeng Liu et al.EMNLP 2025
- Encoding Spreadsheets for Large Language ModelsHaoyu Dong, Jianbo Zhao, Yuzhang Tian, Junyu Xiong et al.EMNLP 2024 · 4 citations
- Semantic table structure identification in spreadsheetsYakun Zhang, Xiao Lv, Haoyu Dong, Wensheng Dou et al.ISSTA 2021 · 11 citations
