Self-Tuning for Data-Efficient Deep Learning
Ximei Wang, Jinghan Gao, Mingsheng Long, Jianmin Wang
Abstract
Deep learning has made revolutionary advances to diverse applications in the presence of large-scale labeled datasets. However, it is prohibitively time-costly and labor-expensive to collect sufficient labeled data in most realistic scenarios. To mitigate the requirement for labeled data, semi-supervised learning (SSL) focuses on simultaneously exploring both labeled and unlabeled data, while transfer learning (TL) popularizes a favorable practice of fine-tuning a pre-trained model to the target data. A dilemma is thus encountered: Without a decent pre-trained model to provide an implicit regularization, SSL through self-training from scratch will be easily misled by inaccurate pseudo-labels, especially in large-sized label space; Without exploring the intrinsic structure of unlabeled data, TL through fine-tuning from limited labeled data is at risk of under-transfer caused by model shift. To escape from this dilemma, we present Self-Tuning to enable data-efficient deep learning by unifying the exploration of labeled and unlabeled data and the transfer of a pre-trained model, as well as a Pseudo Group Contrast (PGC) mechanism to mitigate the reliance on pseudo-labels and boost the tolerance to false labels. Self-Tuning outperforms its SSL and TL counterparts on five tasks by sharp margins, e.g. it doubles the accuracy of fine-tuning on Cars with 15% labels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24868308-5871-4031-8e02-371f11eb19cbCited by top-tier papers20
- Debiased Self-Training for Semi-Supervised LearningBaixu Chen, Junguang Jiang, Ximei Wang, Pengfei Wan et al.NeurIPS 2022 · 162 citations
- Visual Recognition with Deep Nearest CentroidsWenguan Wang, Cheng Han, Tianfei Zhou, Dongfang LiuICLR 2023 · 45 citations
- Don't Stop Pretraining? Make Prompt-based Fine-tuning Powerful LearnerZhengxiang Shi, Aldo LipaniNeurIPS 2023 · 36 citations
- CYCLE: Learning to Self-Refine the Code GenerationYangruibo Ding, Marcus J. Min, Gail E. Kaiser, Baishakhi RayOOPSLA 2024 · 35 citations
- Prototype-Guided Pseudo Labeling for Semi-Supervised Text ClassificationWeiyi Yang, Richong Zhang, Junfan Chen, Lihong Wang et al.ACL 2023 · 28 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
Related papers
- Revisiting Semi-Supervised Learning in the Era of Foundation ModelsPing Zhang, Zheda Mai, Quang-Huy Nguyen, Wei-Lun ChaoNeurIPS 2025 · 9 citations
- Selectivity Drives Productivity: Efficient Dataset Pruning for Enhanced Transfer LearningYihua Zhang, Yimeng Zhang, Aochuan Chen, Jinghan Jia et al.NeurIPS 2023 · 18 citations
- A Realistic Evaluation of Semi-Supervised Learning for Fine-Grained ClassificationJong-Chyi Su, Zezhou Cheng, Subhransu MajiCVPR 2021
- FATE: A Prompt-Tuning-Based Semi-Supervised Learning Framework for Extremely Limited Labeled DataHezhao Liu, Yang Lu, Mengke Li, Yiqun Zhang et al.ACM MM 2025 · 1 citation
- Complementary Benefits of Contrastive Learning and Self-Training Under Distribution ShiftSaurabh Garg, Amrith Setlur, Zachary C. Lipton, Sivaraman Balakrishnan et al.NeurIPS 2023 · 13 citations
