Table Search Using a Deep Contextualized Language Model
Zhiyu Chen, Mohamed Trabelsi, Jeff Heflin, Yinan Xu, Brian D. Davison
Abstract
Pretrained contextualized language models such as BERT have achieved impressive results on various natural language processing benchmarks. Benefiting from multiple pretraining tasks and large scale training corpora, pretrained models can capture complex syntactic word relations. In this paper, we use the deep contextualized language model BERT for the task of ad hoc table retrieval. We investigate how to encode table content considering the table structure and input length limit of BERT. We also propose an approach that incorporates features from prior literature on table retrieval and jointly trains them with BERT. In experiments on public datasets, we show that our best approach can outperform the previous state-of-the-art method and BERT baselines with a large margin under different evaluation metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a55b84c6-a2d0-4157-8a78-65648fbc5a43Cited by top-tier papers14
- Pretrained Generalized Autoregressive Model with Adaptive Probabilistic Label Clusters for Extreme Multi-label Text ClassificationHui Ye, Zhiyu Chen, Da-Han Wang, Brian D. DavisonICML 2020 · 57 citations
- DeepJoin: Joinable Table Discovery with Pre-trained Language ModelsYuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto et al.VLDB 2023 · 53 citations
- StruBERT: Structure-aware BERT for Table Search and MatchingMohamed Trabelsi, Zhiyu Chen, Shuo Zhang, Brian D. Davison et al.WWW 2022 · 52 citations
- Retrieving Complex Tables with Multi-Granular Graph Representation LearningFei Wang, Kexuan Sun, Muhao Chen, Jay Pujara et al.SIGIR 2021 · 34 citations
- Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised ApproachQiming Wang, Raul Castro FernandezSIGMOD 2024 · 20 citations
Related papers
- Making Pre-trained Language Models Great on Tabular PredictionJiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu et al.ICLR 2024 · 72 citations
- TabEmb: Joint Semantic-Structure Embedding for Table AnnotationEhsan Hoseinzade, Ke Wang, Anandharaju Durai RajuACL 2026
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- B-PROP: Bootstrapped Pre-training with Representative Words Prediction for Ad-hoc RetrievalXinyu Ma, Jiafeng Guo, Ruqing Zhang, Yixing Fan et al.SIGIR 2021 · 36 citations
- Webformer: Pre-training with Web Pages for Information RetrievalYu Guo, Zhengyi Ma, Jiaxin Mao, Hongjin Qian et al.SIGIR 2022 · 30 citations
