Data driven semi-supervised learning
Maria-Florina Balcan, Dravyansh Sharma
摘要
We consider a novel data driven approach for designing learning algorithms that can effectively learn with only a small number of labeled examples. This is crucial for modern machine learning applications where labels are scarce or expensive to obtain. We focus on graph-based techniques, where the unlabeled examples are connected in a graph under the implicit assumption that similar nodes likely have similar labels. Over the past decades, several elegant graph-based semi-supervised learning algorithms for how to infer the labels of the unlabeled examples given the graph and a few labeled examples have been proposed. However, the problem of how to create the graph (which impacts the practical usefulness of these methods significantly) has been relegated to domain-specific art and heuristics and no general principles have been proposed. In this work we present a novel data driven approach for learning the graph and provide strong formal guarantees in both the distributional and online learning formalizations. We show how to leverage problem instances coming from an underlying problem domain to learn the graph hyperparameters from commonly used parametric families of graphs that perform well on new instances coming from the same domain. We obtain low regret and efficient algorithms in the online setting, and generalization guarantees in the distributional setting. We also show how to combine several very different similarity metrics and learn multiple hyperparameters, providing general techniques to apply to large classes of problems. We expect some of the tools and techniques we develop along the way to be of interest beyond semi-supervised learning, for data driven algorithms for combinatorial problems more generally.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Provably tuning the ElasticNet across instancesMaria-Florina Balcan, Misha Khodak, Dravyansh Sharma, Ameet TalwalkarNeurIPS 2022 · 被引用 28 次
- Sample complexity of data-driven tuning of model hyperparameters in neural networks with structured parameter-dependent dual functionMaria-Florina Balcan, Anh Nguyen, Dravyansh SharmaNeurIPS 2025 · 被引用 14 次
- No Internal Regret with Non-convex Loss FunctionsDravyansh SharmaAAAI 2024 · 被引用 10 次
- Offline-to-Online Hyperparameter Transfer for Stochastic BanditsDravyansh Sharma, Arun SuggalaAAAI 2025 · 被引用 8 次
- TRiCo: Triadic Game-Theoretic Co-Training for Robust Semi-Supervised LearningHongyang He, Xinyuan Song, Yangfan He, Zeyu Zhang 等NeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper1
相关 Paper
- A Graph-Theoretic Framework for Understanding Open-World Semi-Supervised LearningYiyou Sun, Zhenmei Shi, Yixuan LiNeurIPS 2023 · 被引用 38 次
- Semi-Supervised Metric Learning: A Deep ResurrectionUjjal Kr Dutta, Mehrtash Harandi, Chellu Chandra SekharAAAI 2021 · 被引用 7 次
- DualGraph: Improving Semi-supervised Graph Classification via Dual Contrastive LearningXiao Luo, Wei Ju, Meng Qu, Chong Chen 等ICDE 2022 · 被引用 44 次
- Contrastive and Generative Graph Convolutional Networks for Graph-based Semi-Supervised LearningSheng Wan, Shirui Pan, Jian Yang, Chen GongAAAI 2021 · 被引用 162 次
- Optimal Block-wise Asymmetric Graph Construction for Graph-based Semi-supervised LearningZixing Song, Yifei Zhang, Irwin KingNeurIPS 2023 · 被引用 9 次
