SubTab: Subsetting Features of Tabular Data for Self-Supervised Representation Learning
Talip Ucar, Ehsan Hajiramezanali, Lindsay Edwards
Abstract
Self-supervised learning has been shown to be very effective in learning useful representations, and yet much of the success is achieved in data types such as images, audio, and text. The success is mainly enabled by taking advantage of spatial, temporal, or semantic structure in the data through augmentation. However, such structure may not exist in tabular datasets commonly used in fields such as healthcare, making it difficult to design an effective augmentation method, and hindering a similar progress in tabular data setting. In this paper, we introduce a new framework, Subsetting features of Tabular data (SubTab), that turns the task of learning from tabular data into a multi-view representation learning problem by dividing the input features to multiple subsets. We argue that reconstructing the data from the subset of its features rather than its corrupted version in an autoencoder setting can better capture its underlying latent representation. In this framework, the joint representation can be expressed as the aggregate of latent variables of the subsets at test time, which we refer to as collaborative inference. Our experiments show that the SubTab achieves the state of the art (SOTA) performance of 98.31% on MNIST in tabular setting, on par with CNN-based SOTA models, and surpasses existing baselines on three other real-world datasets by a significant margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c2f0225-3968-4179-97ff-a819453404bfCited by top-tier papers33
- TransTab: Learning Transferable Tabular Transformers Across TablesZifeng Wang, Jimeng SunNeurIPS 2022 · 242 citations
- XTab: Cross-table Pretraining for Tabular TransformersBingzhao Zhu, Xingjian Shi, Nick Erickson, Mu Li et al.ICML 2023 · 111 citations
- Large Language Models Can Automatically Engineer Features for Few-Shot Tabular LearningSungwon Han, Jinsung Yoon, Sercan Ö. Arik, Tomas PfisterICML 2024 · 81 citations
- Optimized Feature Generation for Tabular Data via LLMs with Decision Tree ReasoningJaehyun Nam, Kyuyoung Kim, Seunghyuk Oh, Jihoon Tack et al.NeurIPS 2024 · 78 citations
- Demystifying DeFi MEV Activities in Flashbots BundleZihao Li, Jianfeng Li, Zheyuan He, Xiapu Luo et al.CCS 2023 · 29 citations
Builds on8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 956 citations
Related papers
- SwitchTab: Switched Autoencoders Are Effective Tabular LearnersJing Wu, Suiyao Chen, Qi Zhao, Renat Sergazinov et al.AAAI 2024 · 60 citations
- Representation Space Augmentation for Effective Self-Supervised Learning on Tabular DataMoonjung Eo, Kyungeun Lee, Hye-Seung Cho, Dongmin Kim et al.AAAI 2025 · 2 citations
- VIME: Extending the Success of Self- and Semi-supervised Learning to Tabular DomainJinsung Yoon, Yao Zhang, James Jordon, Mihaela van der SchaarNeurIPS 2020 · 370 citations
- T-JEPA: Augmentation-Free Self-Supervised Learning for Tabular DataHugo Thimonier, José Lucas De Melo Costa, Fabrice Popineau, Arpad Rimmel et al.ICLR 2025
- Auto-GAN: Self-Supervised Collaborative Learning for Medical Image SynthesisBing Cao, Han Zhang, Nannan Wang, Xinbo Gao et al.AAAI 2020 · 94 citations
