Random Matrix Analysis to Balance between Supervised and Unsupervised Learning under the Low Density Separation Assumption
Vasilii Feofanov, Malik Tiomoko, Aladin Virmaux
摘要
We propose a theoretical framework to analyze semi-supervised classification under the low density separation assumption in a high-dimensional regime. In particular, we introduce QLDS, a linear classification model, where the low density separation assumption is implemented via quadratic margin maximization. The algorithm has an explicit solution with rich theoretical properties, and we show that particular cases of our algorithm are the least-square support vector machine in the supervised case, the spectral clustering in the fully unsupervised regime, and a class of semi-supervised graph-based approaches. As such, QLDS establishes a smooth bridge between these supervised and unsupervised learning methods. Using recent advances in the random matrix theory, we formally derive a theoretical evaluation of the classification error in the asymptotic regime. As an application, we derive a hyperparameter selection policy that finds the best balance between the supervised and the unsupervised terms of our learning criterion. Finally, we provide extensive illustrations of our framework, as well as an experimental study on several benchmarks to demonstrate that QLDS, while being computationally more efficient, improves over cross-validation for hyperparameter selection, indicating a high promise of the usage of random matrix theory for semi-supervised model selection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Analysing Multi-Task Regression via Random Matrix Theory with Application to Time Series ForecastingRomain Ilbert, Malik Tiomoko, Cosme Louart, Ambroise Odonnat 等NeurIPS 2024 · 被引用 11 次
- MaNo: Exploiting Matrix Norm for Unsupervised Accuracy Estimation Under Distribution ShiftsRenchunzi Xie, Ambroise Odonnat, Vasilii Feofanov, Weijian Deng 等NeurIPS 2024 · 被引用 10 次
- Evaluating multiple models using labeled and unlabeled dataDivya Shanmugam, Shuvom Sadhuka, Manish Raghavan, John V. Guttag 等NeurIPS 2025 · 被引用 9 次
- Towards Realistic Model Selection for Semi-supervised LearningMuyang Li, Xiaobo Xia, Runze Wu, Fengming Huang 等ICML 2024 · 被引用 2 次
- CLID-MU: Cross-Layer Information Divergence Based Meta Update Strategy for Learning with Noisy LabelsRuofan Hu, Dongyu Zhang, Huayi Zhang, Elke A. RundensteinerKDD 2025
它引用的顶会 Paper13
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2022 · 被引用 173 次
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt 等NeurIPS 2021 · 被引用 170 次
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 被引用 163 次
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy RegimeHugo Cui, Bruno Loureiro, Florent Krzakala, Lenka ZdeborováNeurIPS 2021 · 被引用 109 次
- Poisson Learning: Graph Based Semi-Supervised Learning At Very Low Label RatesJeff Calder, Brendan Cook, Matthew Thorpe, Dejan SlepcevICML 2020 · 被引用 101 次
相关 Paper
- Semi-Supervised Sparse Gaussian Classification: Provable Benefits of Unlabeled DataEyar Azar, Boaz NadlerNeurIPS 2024 · 被引用 5 次
- Quantum Spectral Clustering of Mixed GraphsDaniel Volya, Prabhat MishraDAC 2021 · 被引用 12 次
- On hyperparameter tuning in general clustering problemsmXinjie Fan, Yuguang Yue, Purnamrita Sarkar, Y. X. Rachel WangICML 2020 · 被引用 19 次
- A Graph-Theoretic Framework for Understanding Open-World Semi-Supervised LearningYiyou Sun, Zhenmei Shi, Yixuan LiNeurIPS 2023 · 被引用 38 次
- Sampling-based sublinear low-rank matrix arithmetic framework for dequantizing quantum machine learningNai-Hui Chia, András Gilyén, Tongyang Li, Han-Hsuan Lin 等STOC 2020 · 被引用 105 次
