SheetPT: Spreadsheet Pre-training Based on Hierarchical Attention Network
Ran Jia, Qiyu Li, Zihan Xu, Xiaoyuan Jin, Lun Du, Haoyu Dong, Xiao Lv, Shi Han, Dongmei Zhang
摘要
Spreadsheets are an important and unique type of business document for data storage, analysis and presentation. The distinction between spreadsheets and most other types of digital documents lies in that spreadsheets provide users with high flexibility of data organization on the grid. Existing related techniques mainly focus on the tabular data and are incompetent in understanding the entire sheet. On the one hand, spreadsheets have no explicit separation across tabular data and other information, leaving a gap for the deployment of such techniques. On the other hand, pervasive data dependence and semantic relations across the sheet require comprehensive modeling of all the information rather than only the tables. In this paper, we propose SheetPT, the first pre-training technique on spreadsheets to enable effective representation learning under this scenario. For computational effectiveness and efficiency, we propose the coherent chunk, an intermediate semantic unit of sheet structure; and we accordingly devise a hierarchical attention-based architecture to capture contextual information across different structural granularities. Three pre-training objectives are also designed to ensure sufficient training against millions of spreadsheets. Two representative downstream tasks, formula prediction and sheet structure recognition are utilized to evaluate its capability and the prominent results reveal its superiority over existing state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi 等ICLR 2022 · 被引用 347 次
- TUTA: Tree-based Transformers for Generally Structured Table Pre-trainingZhiruo Wang, Haoyu Dong, Ran Jia, Jia Li 等KDD 2021 · 被引用 88 次
- SpreadsheetCoder: Formula Prediction from Semi-structured ContextXinyun Chen, Petros Maniatis, Rishabh Singh, Charles Sutton 等ICML 2021 · 被引用 63 次
- TabularNet: A Neural Network Architecture for Understanding Semantic Structures of Tabular DataLun Du, Fei Gao, Xu Chen, Ran Jia 等KDD 2021 · 被引用 58 次
相关 Paper
- GetPt: Graph-enhanced General Table Pre-training with Alternate Attention NetworkRan Jia, Haoming Guo, Xiaoyuan Jin, Chao Yan 等KDD 2023 · 被引用 3 次
- FORTAP: Using Formulas for Numerical-Reasoning-Aware Table PretrainingZhoujun Cheng, Haoyu Dong, Ran Jia, Pengfei Wu 等ACL 2022
- SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based ReflectionQin Chen, Yuanyi Ren, Xiaojun Ma, Mugeng Liu 等EMNLP 2025
- Encoding Spreadsheets for Large Language ModelsHaoyu Dong, Jianbo Zhao, Yuzhang Tian, Junyu Xiong 等EMNLP 2024 · 被引用 4 次
- Semantic table structure identification in spreadsheetsYakun Zhang, Xiao Lv, Haoyu Dong, Wensheng Dou 等ISSTA 2021 · 被引用 11 次
