TUTA: Tree-based Transformers for Generally Structured Table Pre-training
Zhiruo Wang, Haoyu Dong, Ran Jia, Jia Li, Zhiyi Fu, Shi Han, Dongmei Zhang
摘要
Tables are widely used with various structures to organize and present data. Recent attempts on table understanding mainly focus on relational tables, yet overlook to other common table structures. In this paper, we propose TUTA, a unified pre-training architecture for understanding generally structured tables. Noticing that understanding a table requires spatial, hierarchical, and semantic information, we enhance transformers with three novel structureaware mechanisms. First, we devise a unified tree-based structure, called a bi-dimensional coordinate tree, to describe both the spatial and hierarchical information of generally structured tables. Upon this, we propose tree-based attention and position embedding to better capture the spatial and hierarchical information. Moreover, we devise three progressive pre-training objectives to enable representations at the token, cell, and table levels. We pre-train TUTA on a wide range of unlabeled web and spreadsheet tables and fine-tune it on two critical tasks in the field of table structure understanding: cell type classification and table type classification. Experiments show that TUTA is highly effective, achieving state-of-the-art on five widely-studied datasets. CCS CONCEPTS • Information systems → Information retrieval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper51
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi 等ICLR 2022 · 被引用 347 次
- Chain-of-Table: Evolving Tables in the Reasoning Chain for Table UnderstandingZilong Wang, Hao Zhang, Chun-Liang Li, Julian Martin Eisenschlos 等ICLR 2024 · 被引用 244 次
- MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual DataYilun Zhao, Yunxiang Li, Chenying Li, Rui ZhangACL 2022 · 被引用 168 次
- TableFormer: Robust Transformer Modeling for Table-Text EncodingJingfeng Yang, Aditya Gupta, Shyam Upadhyay, Luheng He 等ACL 2022 · 被引用 145 次
- Annotating Columns with Pre-trained Language ModelsYoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang 等SIGMOD 2022 · 被引用 81 次
它引用的顶会 Paper8
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- Tree-Structured Attention with Hierarchical AccumulationXuan-Phi Nguyen, Shafiq R. Joty, Steven C. H. Hoi, Richard SocherICLR 2020 · 被引用 79 次
- Table2Analysis: Modeling and Recommendation of Common Analysis Patterns for Multi-Dimensional DataMengyu Zhou, Wang Tao, Pengxin Ji, Han Shi 等AAAI 2020 · 被引用 26 次
相关 Paper
- GetPt: Graph-enhanced General Table Pre-training with Alternate Attention NetworkRan Jia, Haoming Guo, Xiaoyuan Jin, Chao Yan 等KDD 2023 · 被引用 3 次
- FORTAP: Using Formulas for Numerical-Reasoning-Aware Table PretrainingZhoujun Cheng, Haoyu Dong, Ran Jia, Pengfei Wu 等ACL 2022
- TabularNet: A Neural Network Architecture for Understanding Semantic Structures of Tabular DataLun Du, Fei Gao, Xu Chen, Ran Jia 等KDD 2021 · 被引用 58 次
- TableVLM: Multi-modal Pre-training for Table Structure RecognitionLeiyuan Chen, Chengsong Huang, Xiaoqing Zheng, Jinshu Lin 等ACL 2023 · 被引用 8 次
- Large Language Models are Complex Table ParsersBowen Zhao, Changkai Ji, Yuejie Zhang, Wen He 等EMNLP 2023 · 被引用 17 次
