UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data Science
Yazheng Yang, Yuqi Wang, Guang Liu, Ledell Wu, Qi Liu
摘要
Recent advancements in Natural Language Processing (NLP) have witnessed the groundbreaking impact of pretrained models, yielding impressive outcomes across various tasks. This study seeks to extend the power of pretraining methodologies to facilitating the prediction over tables in data science, a domain traditionally overlooked, yet inherently challenging due to the plethora of table schemas intrinsic to different tasks. The primary research questions underpinning this work revolve around the establishment of a universal pretraining protocol for tables with varied structures, the generalizability and transferability of learned knowledge across tasks, the adaptation to diverse downstream applications, and the incorporation of incremental columns over time. In response to these challenges, we introduce UniTabE, a straightforward yet effective method designed to process tables in a uniform manner, devoid of constraints imposed by specific table structures. UniTabE's core concept relies on representing each basic table element with a module, termed TabUnit. This is subsequently followed by a Transformer encoder to refine the representation. Moreover, our model is designed to facilitate pretraining and finetuning through the utilization of free-form prompts. In order to implement the pretraining phase, we curated an expansive tabular dataset comprising approximately 13 billion samples, meticulously gathered from the Kaggle platform. This research primarily centers on classification and regression tasks involving tabular data, and conducts rigorous experimental testing and analyses to validate the effectiveness of our methodology. The experimental results demonstrate UniTabE's superior performance against several baseline models across a multitude of benchmark datasets. This, therefore, underscores UniTabE's potential to significantly enhance the semantic representation of tabular data, thereby marking a significant stride for tabular data analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation ModelsXiyuan Zhang, Danielle Maddix Robinson, Junming Yin, Nick Erickson 等NeurIPS 2025 · 被引用 91 次
- From Supervised to Generative: A Novel Paradigm for Tabular Deep Learning with Large Language ModelsXumeng Wen, Han Zhang, Shun Zheng, Wei Xu 等KDD 2024 · 被引用 9 次
- Handling Learnwares from Heterogeneous Feature Spaces with Explicit Label ExploitationPeng Tan, Hai-Tian Liu, Zhi-Hao Tan, Zhi-Hua ZhouNeurIPS 2024 · 被引用 8 次
- Griffin: Towards a Graph-Centric Relational Database Foundation ModelYanbo Wang, Xiyuan Wang, Quan Gan, Minjie Wang 等ICML 2025
- Robust Detection of Synthetic Tabular Data Under Schema VariabilityG. Charbel N. Kindji, Elisa Fromont, Lina Maria Rojas-Barahona, Tanguy UrvoyAAAI 2026
它引用的顶会 Paper14
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 被引用 2,148 次
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 被引用 1,847 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular DataSergei Popov, Stanislav Morozov, Artem BabenkoICLR 2020 · 被引用 407 次
相关 Paper
- Towards Cross-Table Masked Pretraining for Web Data MiningChao Ye, Guoshan Lu, Haobo Wang, Liyao Li 等WWW 2024 · 被引用 23 次
- TransTab: Learning Transferable Tabular Transformers Across TablesZifeng Wang, Jimeng SunNeurIPS 2022 · 被引用 242 次
- XTab: Cross-table Pretraining for Tabular TransformersBingzhao Zhu, Xingjian Shi, Nick Erickson, Mu Li 等ICML 2023 · 被引用 111 次
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- GetPt: Graph-enhanced General Table Pre-training with Alternate Attention NetworkRan Jia, Haoming Guo, Xiaoyuan Jin, Chao Yan 等KDD 2023 · 被引用 3 次
