End-to-End Weak Supervision
Salva Rühling Cachay, Benedikt Boecking, Artur Dubrawski
摘要
Aggregating multiple sources of weak supervision (WS) can ease the data-labeling bottleneck prevalent in many machine learning applications, by replacing the tedious manual collection of ground truth labels. Current state of the art approaches that do not use any labeled training data, however, require two separate modeling steps: Learning a probabilistic latent variable model based on the WS sources -- making assumptions that rarely hold in practice -- followed by downstream model training. Importantly, the first step of modeling does not consider the performance of the downstream model. To address these caveats we propose an end-to-end approach for directly learning the downstream model by maximizing its agreement with probabilistic labels generated by reparameterizing previous probabilistic posteriors with a neural network. Our results show improved performance over prior work in terms of end model performance on downstream test sets, as well as in terms of improved robustness to dependencies among weak supervision sources.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Theoretical Analysis of Weak-to-Strong GeneralizationHunter Lang, David A. Sontag, Aravindan VijayaraghavanNeurIPS 2024 · 被引用 59 次
- Nemo: Guiding and Contextualizing Weak Supervision for Interactive Data ProgrammingCheng-Yu Hsieh, Jieyu Zhang, Alexander J. RatnerVLDB 2022 · 被引用 17 次
- Losses over Labels: Weakly Supervised Learning via Direct Loss ConstructionDylan Sam, J. Zico KolterAAAI 2023 · 被引用 14 次
- Characterizing the Impacts of Semi-supervised Learning for Weak SupervisionJeffrey Li, Jieyu Zhang, Ludwig Schmidt, Alexander J. RatnerNeurIPS 2023 · 被引用 9 次
- CARE: Confounder-Aware Aggregation for Reliable LLM EvaluationJitian Zhao, Changho Shin, Tzu-Heng Huang, Satya Sai Srinath Namburi GNVV 等ICML 2026 · 被引用 7 次
它引用的顶会 Paper6
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Understanding self-supervised learning dynamics without contrastive pairsYuandong Tian, Xinlei Chen, Surya GanguliICML 2021 · 被引用 338 次
- Fast and Three-rious: Speeding Up Weak Supervision with Triplet MethodsDaniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper 等ICML 2020 · 被引用 130 次
- Learning from Rules Generalizing Labeled ExemplarsAbhijeet Awasthi, Sabyasachi Ghosh, Rasna Goyal, Sunita SarawagiICLR 2020 · 被引用 93 次
- Interactive Weak Supervision: Learning Useful Heuristics for Data LabelingBenedikt Boecking, Willie Neiswanger, Eric P. Xing, Artur DubrawskiICLR 2021 · 被引用 8 次
相关 Paper
- Generative Modeling Helps Weak Supervision (and Vice Versa)Benedikt Boecking, Nicholas Carl Roberts, Willie Neiswanger, Stefano Ermon 等ICLR 2023 · 被引用 1 次
- Learning Hyper Label Model for Programmatic Weak SupervisionRenzhi Wu, Shen-En Chen, Jieyu Zhang, Xu ChuICLR 2023 · 被引用 2 次
- Creating Training Sets via Weak Indirect SupervisionJieyu Zhang, Bohan Wang, Xiangchen Song, Yujing Wang 等ICLR 2022 · 被引用 17 次
- Understanding Programmatic Weak Supervision via Source-aware Influence FunctionJieyu Zhang, Haonan Wang, Cheng-Yu Hsieh, Alexander J. RatnerNeurIPS 2022 · 被引用 13 次
- Amortized Variational Inference for Partial-Label Learning: A Probabilistic Approach to Label DisambiguationTobias Fuchs, Nadja KleinICML 2026
