VIME: Extending the Success of Self- and Semi-supervised Learning to Tabular Domain
Jinsung Yoon, Yao Zhang, James Jordon, Mihaela van der Schaar
摘要
Self-and semi-supervised learning frameworks have made significant progress in training machine learning models with limited labeled data in image and language domains. These methods heavily rely on the unique structure in the domain datasets (such as spatial relationships in images or semantic relationships in language). They are not adaptable to general tabular data which does not have the same explicit structure as image and language data. In this paper, we fill this gap by proposing novel self-and semi-supervised learning frameworks for tabular data, which we refer to collectively as VIME (Value Imputation and Mask Estimation). We create a novel pretext task of estimating mask vectors from corrupted tabular data in addition to the reconstruction pretext task for self-supervised learning. We also introduce a novel tabular data augmentation method for self-and semi-supervised learning frameworks. In experiments, we evaluate the proposed framework in multiple tabular datasets from various application domains, such as genomics and clinical data. VIME exceeds state-of-the-art performance in comparison to the existing baseline methods. Unfortunately, existing self-and semi-supervised learning algorithms are not effective for tabular data 1 because they heavily rely on the spatial or semantic structure of image or language data. A 1 Tabular data is a database that is structured in a tabular form. It arranges data elements in vertical columns (features) and horizontal rows (samples). 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- TransTab: Learning Transferable Tabular Transformers Across TablesZifeng Wang, Jimeng SunNeurIPS 2022 · 被引用 242 次
- Scarf: Self-Supervised Contrastive Learning using Random Feature CorruptionDara Bahri, Heinrich Jiang, Yi Tay, Donald MetzlerICLR 2022 · 被引用 233 次
- SubTab: Subsetting Features of Tabular Data for Self-Supervised Representation LearningTalip Ucar, Ehsan Hajiramezanali, Lindsay EdwardsNeurIPS 2021 · 被引用 189 次
- TabDPT: Scaling Tabular Foundation Models on Real DataJunwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach 等NeurIPS 2025 · 被引用 118 次
- Optimized Feature Generation for Tabular Data via LLMs with Decision Tree ReasoningJaehyun Nam, Kyuyoung Kim, Seunghyuk Oh, Jihoon Tack 等NeurIPS 2024 · 被引用 78 次
它引用的顶会 Paper4
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 被引用 2,148 次
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 被引用 1,553 次
- Momentum Contrast for Unsupervised Visual Representation LearningKaiming He, Haoqi Fan, Yuxin Wu, Saining Xie 等CVPR 2020
相关 Paper
- STUNT: Few-shot Tabular Learning with Self-generated Tasks from Unlabeled TablesJaehyun Nam, Jihoon Tack, Kyungmin Lee, Hankook Lee 等ICLR 2023 · 被引用 2 次
- T-JEPA: Augmentation-Free Self-Supervised Learning for Tabular DataHugo Thimonier, José Lucas De Melo Costa, Fabrice Popineau, Arpad Rimmel 等ICLR 2025
- Representation Space Augmentation for Effective Self-Supervised Learning on Tabular DataMoonjung Eo, Kyungeun Lee, Hye-Seung Cho, Dongmin Kim 等AAAI 2025 · 被引用 2 次
- SMILE: Infusing Spatial and Motion Semantics in Masked Video LearningFida Mohammad Thoker, Letian Jiang, Chen Zhao, Bernard GhanemCVPR 2025
- SwitchTab: Switched Autoencoders Are Effective Tabular LearnersJing Wu, Suiyao Chen, Qi Zhao, Renat Sergazinov 等AAAI 2024 · 被引用 60 次
