Scarf: Self-Supervised Contrastive Learning using Random Feature Corruption
Dara Bahri, Heinrich Jiang, Yi Tay, Donald Metzler
Abstract
Self-supervised contrastive representation learning has proved incredibly successful in the vision and natural language domains, enabling state-of-the-art performance with orders of magnitude less labeled data. However, such methods are domain-specific and little has been done to leverage this technique on real-world tabular datasets. We propose SCARF, a simple, widely-applicable technique for contrastive learning, where views are formed by corrupting a random subset of features. When applied to pre-train deep neural networks on the 69 real-world, tabular classification datasets from the OpenML-CC18 benchmark, SCARF not only improves classification accuracy in the fully-supervised setting but does so also in the presence of label noise and in the semi-supervised setting where only a fraction of the available training data is labeled. We show that SCARF complements existing strategies and outperforms alternatives like autoencoders. We conduct comprehensive ablations, detailing the importance of a range of factors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de272dd4-93c5-4a04-b010-183cc8e5d096Cited by top-tier papers51
- TransTab: Learning Transferable Tabular Transformers Across TablesZifeng Wang, Jimeng SunNeurIPS 2022 · 242 citations
- TabDPT: Scaling Tabular Foundation Models on Real DataJunwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach et al.NeurIPS 2025 · 118 citations
- XTab: Cross-table Pretraining for Tabular TransformersBingzhao Zhu, Xingjian Shi, Nick Erickson, Mu Li et al.ICML 2023 · 111 citations
- Large Language Models Can Automatically Engineer Features for Few-Shot Tabular LearningSungwon Han, Jinsung Yoon, Sercan Ö. Arik, Tomas PfisterICML 2024 · 81 citations
- Fascinating Supervisory Signals and Where to Find Them: Deep Anomaly Detection with Scale LearningHongzuo Xu, Yijie Wang, Juhui Wei, Songlei Jian et al.ICML 2023 · 65 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
Related papers
- Best of Both Worlds: Multimodal Contrastive Learning with Tabular and Imaging DataPaul Hager, Martin J. Menten, Daniel RueckertCVPR 2023
- Towards Domain-Agnostic Contrastive LearningVikas Verma, Thang Luong, Kenji Kawaguchi, Hieu Pham et al.ICML 2021 · 131 citations
- Investigating Why Contrastive Learning Benefits Robustness against Label NoiseYihao Xue, Kyle Whitecross, Baharan MirzasoleimanICML 2022 · 70 citations
- Adversarial Self-Supervised Contrastive LearningMinseon Kim, Jihoon Tack, Sung Ju HwangNeurIPS 2020 · 294 citations
- Selective-Supervised Contrastive Learning with Noisy LabelsShikun Li, Xiaobo Xia, Shiming Ge, Tongliang LiuCVPR 2022 · 201 citations
