Random Matrix Analysis to Balance between Supervised and Unsupervised Learning under the Low Density Separation Assumption
Vasilii Feofanov, Malik Tiomoko, Aladin Virmaux
Abstract
We propose a theoretical framework to analyze semi-supervised classification under the low density separation assumption in a high-dimensional regime. In particular, we introduce QLDS, a linear classification model, where the low density separation assumption is implemented via quadratic margin maximization. The algorithm has an explicit solution with rich theoretical properties, and we show that particular cases of our algorithm are the least-square support vector machine in the supervised case, the spectral clustering in the fully unsupervised regime, and a class of semi-supervised graph-based approaches. As such, QLDS establishes a smooth bridge between these supervised and unsupervised learning methods. Using recent advances in the random matrix theory, we formally derive a theoretical evaluation of the classification error in the asymptotic regime. As an application, we derive a hyperparameter selection policy that finds the best balance between the supervised and the unsupervised terms of our learning criterion. Finally, we provide extensive illustrations of our framework, as well as an experimental study on several benchmarks to demonstrate that QLDS, while being computationally more efficient, improves over cross-validation for hyperparameter selection, indicating a high promise of the usage of random matrix theory for semi-supervised model selection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ab0a5c7-5197-4c8a-8a22-9bb09be3d27aCited by top-tier papers5
- Analysing Multi-Task Regression via Random Matrix Theory with Application to Time Series ForecastingRomain Ilbert, Malik Tiomoko, Cosme Louart, Ambroise Odonnat et al.NeurIPS 2024 · 11 citations
- MaNo: Exploiting Matrix Norm for Unsupervised Accuracy Estimation Under Distribution ShiftsRenchunzi Xie, Ambroise Odonnat, Vasilii Feofanov, Weijian Deng et al.NeurIPS 2024 · 10 citations
- Evaluating multiple models using labeled and unlabeled dataDivya Shanmugam, Shuvom Sadhuka, Manish Raghavan, John V. Guttag et al.NeurIPS 2025 · 9 citations
- Towards Realistic Model Selection for Semi-supervised LearningMuyang Li, Xiaobo Xia, Runze Wu, Fengming Huang et al.ICML 2024 · 2 citations
- CLID-MU: Cross-Layer Information Divergence Based Meta Update Strategy for Learning with Noisy LabelsRuofan Hu, Dongyu Zhang, Huayi Zhang, Elke A. RundensteinerKDD 2025
Builds on13
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang et al.NeurIPS 2022 · 173 citations
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt et al.NeurIPS 2021 · 170 citations
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 163 citations
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy RegimeHugo Cui, Bruno Loureiro, Florent Krzakala, Lenka ZdeborováNeurIPS 2021 · 109 citations
- Poisson Learning: Graph Based Semi-Supervised Learning At Very Low Label RatesJeff Calder, Brendan Cook, Matthew Thorpe, Dejan SlepcevICML 2020 · 101 citations
Related papers
- Semi-Supervised Sparse Gaussian Classification: Provable Benefits of Unlabeled DataEyar Azar, Boaz NadlerNeurIPS 2024 · 5 citations
- Quantum Spectral Clustering of Mixed GraphsDaniel Volya, Prabhat MishraDAC 2021 · 12 citations
- On hyperparameter tuning in general clustering problemsmXinjie Fan, Yuguang Yue, Purnamrita Sarkar, Y. X. Rachel WangICML 2020 · 19 citations
- A Graph-Theoretic Framework for Understanding Open-World Semi-Supervised LearningYiyou Sun, Zhenmei Shi, Yixuan LiNeurIPS 2023 · 38 citations
- Sampling-based sublinear low-rank matrix arithmetic framework for dequantizing quantum machine learningNai-Hui Chia, András Gilyén, Tongyang Li, Han-Hsuan Lin et al.STOC 2020 · 105 citations
