Numerical Tuple Extraction from Tables with Pre-training
Qingping Yang, Yixuan Cao, Ping Luo
Abstract
Tables are omnipresent on the web and in various vertical domains, storing massive amounts of valuable data. However, the great flexibility in the table layout hinders the machine from understanding this valuable data. In order to unlock and utilize knowledge from tables, extracting data as numerical tuples is the first and critical step. As a form of relational data, numerical tuples have direct and transparent relationships between their elements and are therefore easy for machines to use. Extracting numerical tuples requires a deep understanding of intricate correlations between cells. The correlations are presented implicitly in texts and visual appearances of tables, which can be roughly classified into Hierarchy and Juxtaposition. Although many studies have made considerable progress in data extraction from tables, most of them only consider hierarchical relationships but neglect the juxtapositions. Meanwhile, they only evaluate their methods on relatively small corpora. This paper proposes a new framework to extract numerical tuples from tables and evaluate it on a large test set. Specifically, we convert this task into a relation extraction problem between cells. To represent cells with their intricate correlations in tables, we propose a BERT-based pre-trained language model, TableLM, to encode tables with diverse layouts. To evaluate the framework, we collect a large finance dataset that includes 19,264 tables and 604K tuples. Extensive experiments on the dataset are conducted to demonstrate the superiority of our framework compared to a well-designed baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04dbec17-44ed-406b-ae3b-d4a02f753049Builds on9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- TUTA: Tree-based Transformers for Generally Structured Table Pre-trainingZhiruo Wang, Haoyu Dong, Ran Jia, Jia Li et al.KDD 2021 · 88 citations
Related papers
- Making Pre-trained Language Models Great on Tabular PredictionJiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu et al.ICLR 2024 · 72 citations
- TabularBERT: Binning-Based Self-Supervised Learning for Tabular RepresentationBeomjin Park, Seunghwan An, Sungchul Hong, Hosik ChoiICML 2026
- TCN: Table Convolutional Network for Web Table InterpretationDaheng Wang, Prashant Shiralkar, Colin Lockard, Binxuan Huang et al.WWW 2021 · 68 citations
- FiNER: Financial Numeric Entity Recognition for XBRL TaggingLefteris Loukas, Manos Fergadiotis, Ilias Chalkidis, Eirini Spyropoulou et al.ACL 2022 · 90 citations
- Table Search Using a Deep Contextualized Language ModelZhiyu Chen, Mohamed Trabelsi, Jeff Heflin, Yinan Xu et al.SIGIR 2020 · 48 citations
