Sinkhorn Label Allocation: Semi-Supervised Classification via Annealed Self-Training
Kai Sheng Tai, Peter Bailis, Gregory Valiant
Abstract
Self-training is a standard approach to semi-supervised learning where the learner's own predictions on unlabeled data are used as supervision during training. In this paper, we reinterpret this label assignment process as an optimal transportation problem between examples and classes, wherein the cost of assigning an example to a class is mediated by the current predictions of the classifier. This formulation facilitates a practical annealing strategy for label assignment and allows for the inclusion of prior knowledge on class proportions via flexible upper bound constraints. The solutions to these assignment problems can be efficiently approximated using Sinkhorn iteration, thus enabling their use in the inner loop of standard stochastic optimization algorithms. We demonstrate the effectiveness of our algorithm on the CIFAR-10, CIFAR-100, and SVHN datasets in comparison with FixMatch, a state-of-the-art self-training algorithm. Our code is available at https://github.com/stanford-futuredata/sinkhorn-label-allocation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f03678d0-754e-4bdb-9f2d-c4f91f71bb97Cited by top-tier papers20
- Self-Training Multi-Sequence Learning with Transformer for Weakly Supervised Video Anomaly DetectionShuo Li, Fang Liu, Licheng JiaoAAAI 2022 · 282 citations
- SoLar: Sinkhorn Label Refinery for Imbalanced Partial-Label LearningHaobo Wang, Mingxuan Xia, Yixuan Li, Yuren Mao et al.NeurIPS 2022 · 54 citations
- Smoothed Adaptive Weighting for Imbalanced Semi-Supervised Learning: Improve Reliability Against Unknown Distribution DataZhengfeng Lai, Chao Wang, Henrry Gunawan, Sen-Ching S. Cheung et al.ICML 2022 · 52 citations
- CSOT: Curriculum and Structure-Aware Optimal Transport for Learning with Noisy LabelsWanxing Chang, Ye Shi, Jingya WangNeurIPS 2023 · 24 citations
- Towards Semi-supervised Learning with Non-random Missing LabelsYue Duan, Zhen Zhao, Lei Qi, Luping Zhou et al.ICCV 2023 · 22 citations
Builds on5
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- ReMixMatch: Semi-Supervised Learning with Distribution Matching and Augmentation AnchoringDavid Berthelot, Nicholas Carlini, Ekin D. Cubuk, Alex Kurakin et al.ICLR 2020 · 469 citations
- Self-Training With Noisy Student Improves ImageNet ClassificationQizhe Xie, Minh-Thang Luong, Eduard H. Hovy, Quoc V. LeCVPR 2020
Related papers
- OTMatch: Improving Semi-Supervised Learning with Optimal TransportZhiquan Tan, Kaipeng Zheng, Weiran HuangICML 2024 · 10 citations
- Debiased Self-Training for Semi-Supervised LearningBaixu Chen, Junguang Jiang, Ximei Wang, Pengfei Wan et al.NeurIPS 2022 · 162 citations
- Label Matching Semi-Supervised Object DetectionBinbin Chen, Weijie Chen, Shicai Yang, Yunyi Xuan et al.CVPR 2022 · 87 citations
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 873 citations
- Progressive Distribution Matching for Federated Semi-Supervised LearningDongping Liao, Xitong Gao, Yabo Xu, Cheng-Zhong XuAAAI 2025 · 1 citation
