Table Search Using a Deep Contextualized Language Model
Zhiyu Chen, Mohamed Trabelsi, Jeff Heflin, Yinan Xu, Brian D. Davison
摘要
Pretrained contextualized language models such as BERT have achieved impressive results on various natural language processing benchmarks. Benefiting from multiple pretraining tasks and large scale training corpora, pretrained models can capture complex syntactic word relations. In this paper, we use the deep contextualized language model BERT for the task of ad hoc table retrieval. We investigate how to encode table content considering the table structure and input length limit of BERT. We also propose an approach that incorporates features from prior literature on table retrieval and jointly trains them with BERT. In experiments on public datasets, we show that our best approach can outperform the previous state-of-the-art method and BERT baselines with a large margin under different evaluation metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Pretrained Generalized Autoregressive Model with Adaptive Probabilistic Label Clusters for Extreme Multi-label Text ClassificationHui Ye, Zhiyu Chen, Da-Han Wang, Brian D. DavisonICML 2020 · 被引用 57 次
- DeepJoin: Joinable Table Discovery with Pre-trained Language ModelsYuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto 等VLDB 2023 · 被引用 53 次
- StruBERT: Structure-aware BERT for Table Search and MatchingMohamed Trabelsi, Zhiyu Chen, Shuo Zhang, Brian D. Davison 等WWW 2022 · 被引用 52 次
- Retrieving Complex Tables with Multi-Granular Graph Representation LearningFei Wang, Kexuan Sun, Muhao Chen, Jay Pujara 等SIGIR 2021 · 被引用 34 次
- Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised ApproachQiming Wang, Raul Castro FernandezSIGMOD 2024 · 被引用 20 次
相关 Paper
- Making Pre-trained Language Models Great on Tabular PredictionJiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu 等ICLR 2024 · 被引用 72 次
- TabEmb: Joint Semantic-Structure Embedding for Table AnnotationEhsan Hoseinzade, Ke Wang, Anandharaju Durai RajuACL 2026
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- B-PROP: Bootstrapped Pre-training with Representative Words Prediction for Ad-hoc RetrievalXinyu Ma, Jiafeng Guo, Ruqing Zhang, Yixing Fan 等SIGIR 2021 · 被引用 36 次
- Webformer: Pre-training with Web Pages for Information RetrievalYu Guo, Zhengyi Ma, Jiaxin Mao, Hongjin Qian 等SIGIR 2022 · 被引用 30 次
