CARTE: Pretraining and Transfer for Tabular Learning
Myung Jun Kim, Léo Grinsztajn, Gaël Varoquaux
摘要
Pretrained deep-learning models are the go-to solution for images or text. However, for tabular data the standard is still to train tree-based models. Indeed, transfer learning on tables hits the challenge of data integration: finding correspondences, correspondences in the entries (entity matching) where different words may denote the same entity, correspondences across columns (schema matching), which may come in different orders, names... We propose a neural architecture that does not need such correspondences. As a result, we can pretrain it on background data that has not been matched. The architecture -- CARTE for Context Aware Representation of Table Entries -- uses a graph representation of tabular (or relational) data to process tables with different columns, string embedding of entries and columns names to model an open vocabulary, and a graph-attentional network to contextualize entries with column names and neighboring entries. An extensive benchmark shows that CARTE facilitates learning, outperforming a solid set of baselines including the best tree-based models. CARTE also enables joint learning across tables with unmatched columns, enhancing a small table with bigger ones. CARTE opens the door to large pretrained models for tabular data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Better by default: Strong pre-tuned MLPs and boosted trees on tabular dataDavid Holzmüller, Léo Grinsztajn, Ingo SteinwartNeurIPS 2024 · 被引用 141 次
- TabDPT: Scaling Tabular Foundation Models on Real DataJunwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach 等NeurIPS 2025 · 被引用 118 次
- Large Scale Transfer Learning for Tabular Data via Language ModelingJosh Gardner, Juan C. Perdomo, Ludwig SchmidtNeurIPS 2024 · 被引用 103 次
- Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation ModelsXiyuan Zhang, Danielle Maddix Robinson, Junming Yin, Nick Erickson 等NeurIPS 2025 · 被引用 91 次
- A Closer Look at TabPFN v2: Understanding Its Strengths and Extending Its CapabilitiesHan-Jia Ye, Si-Yang Liu, Wei-Lun ChaoNeurIPS 2025 · 被引用 52 次
它引用的顶会 Paper8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 被引用 2,148 次
- Measuring and Relieving the Over-Smoothing Problem for Graph Neural Networks from the Topological ViewDeli Chen, Yankai Lin, Wei Li, Peng Li 等AAAI 2020 · 被引用 1,353 次
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular DataSergei Popov, Stanislav Morozov, Artem BabenkoICLR 2020 · 被引用 407 次
相关 Paper
- ConTextTab: A Semantics-Aware Tabular In-Context LearnerMarco Spinaci, Marek Polewczyk, Maximilian Schambach, Sam ThelinNeurIPS 2025 · 被引用 36 次
- GetPt: Graph-enhanced General Table Pre-training with Alternate Attention NetworkRan Jia, Haoming Guo, Xiaoyuan Jin, Chao Yan 等KDD 2023 · 被引用 3 次
- TCN: Table Convolutional Network for Web Table InterpretationDaheng Wang, Prashant Shiralkar, Colin Lockard, Binxuan Huang 等WWW 2021 · 被引用 68 次
- TransTab: Learning Transferable Tabular Transformers Across TablesZifeng Wang, Jimeng SunNeurIPS 2022 · 被引用 242 次
- Making Pre-trained Language Models Great on Tabular PredictionJiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu 等ICLR 2024 · 被引用 72 次
