Lifting Weak Supervision To Structured Prediction
Harit Vishwakarma, Frederic Sala
摘要
Weak supervision (WS) is a rich set of techniques that produce pseudolabels by aggregating easily obtained but potentially noisy label estimates from a variety of sources. WS is theoretically well understood for binary classification, where simple approaches enable consistent estimation of pseudolabel noise rates. Using this result, it has been shown that downstream models trained on the pseudolabels have generalization guarantees nearly identical to those trained on clean labels. While this is exciting, users often wish to use WS for structured prediction, where the output space consists of more than a binary or multi-class label set: e.g. rankings, graphs, manifolds, and more. Do the favorable theoretical properties of WS for binary classification lift to this setting? We answer this question in the affirmative for a wide range of scenarios. For labels taking values in a finite metric space, we introduce techniques new to weak supervision based on pseudo-Euclidean embeddings and tensor decompositions, providing a nearly-consistent noise rate estimator. For labels in constant-curvature Riemannian manifolds, we introduce new invariants that also yield consistent noise rate estimation. In both cases, when using the resulting pseudolabels in concert with a flexible downstream model, we obtain generalization guarantees nearly identical to those for models trained on clean data. Several of our results, which can be viewed as robustness guarantees in structured prediction with noisy labels, may be of independent interest. Empirical evaluation validates our claims and shows the merits of the proposed method 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Smoothie: Label Free Language Model RoutingNeel Guha, Mayee F. Chen, Trevor Chow, Ishan S. Khare 等NeurIPS 2024 · 被引用 44 次
- Promises and Pitfalls of Threshold-based Auto-labelingHarit Vishwakarma, Heguang Lin, Frederic Sala, Ramya Korlakai VinayakNeurIPS 2023 · 被引用 16 次
- Pearls from Pebbles: Improved Confidence Functions for Auto-labelingHarit Vishwakarma, Yi Chen, Sui Jiet Tay, Satya Sai Srinath Namburi 等NeurIPS 2024 · 被引用 7 次
- Embroid: Unsupervised Prediction Smoothing Can Improve Few-Shot ClassificationNeel Guha, Mayee F. Chen, Kush Bhatia, Azalia Mirhoseini 等NeurIPS 2023 · 被引用 6 次
- Weaver: Shrinking the Generation-Verification Gap by Scaling Compute for VerificationJon Saad-Falcon, Estefany Kelly Buchanan, Mayee F. Chen, Tzu-Heng Huang 等NeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper2
相关 Paper
- Universalizing Weak SupervisionChangho Shin, Winfred Li, Harit Vishwakarma, Nicholas Carl Roberts 等ICLR 2022 · 被引用 35 次
- Creating Training Sets via Weak Indirect SupervisionJieyu Zhang, Bohan Wang, Xiangchen Song, Yujing Wang 等ICLR 2022 · 被引用 17 次
- Limited-Supervised Multi-Label Learning with Dependency NoiseYejiang Wang, Yuhai Zhao, Zhengkui Wang, Wen Shan 等AAAI 2024 · 被引用 7 次
- Learning Hyper Label Model for Programmatic Weak SupervisionRenzhi Wu, Shen-En Chen, Jieyu Zhang, Xu ChuICLR 2023 · 被引用 2 次
- Task Agnostic Robust Learning on Corrupt Outputs by Correlation-Guided Mixture Density NetworksSungjoon Choi, Sanghoon Hong, Kyungjae Lee, Sungbin LimCVPR 2020
