Foreseeing the Benefits of Incidental Supervision
Hangfeng He, Mingyuan Zhang, Qiang Ning, Dan Roth
Abstract
Real-world applications often require improved models by leveraging a range of cheap incidental supervision signals. These could include partial labels, noisy labels, knowledgebased constraints, and cross-domain or crosstask annotations -all having statistical associations with gold annotations but not exactly the same. However, we currently lack a principled way to measure the benefits of these signals to a given target task, and the common practice of evaluating these benefits is through exhaustive experiments with various models and hyperparameters. This paper studies whether we can, in a single framework, quantify the benefits of various types of incidental signals for a given target task without going through combinatorial experiments. We propose a unified PAC-Bayesian motivated informativeness measure, PABI, that characterizes the uncertainty reduction provided by incidental supervision signals. We demonstrate PABI's effectiveness by quantifying the value added by various types of incidental signals to sequence tagging tasks. Experiments on named entity recognition (NER) and question answering (QA) show that PABI's predictions correlate well with learning performance, providing a promising way to determine, ahead of learning, which supervision signals would be beneficial. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Weighted Training for Cross-Task LearningShuxiao Chen, Koby Crammer, Hangfeng He, Dan Roth et al.ICLR 2022 · 30 citations
- Can NLI Provide Proper Indirect Supervision for Low-resource Biomedical Relation Extraction?Jiashu Xu, Mingyu Derek Ma, Muhao ChenACL 2023 · 14 citations
- Surveying the Dead Minds: Historical-Psychological Text Analysis with Contextualized Construct Representation (CCR) for Classical ChineseYuqi Chen, Sixuan Li, Ying Li, Mohammad AtariEMNLP 2024 · 10 citations
Builds on4
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Learning from Noisy Labels with No Change to the Training ProcessMingyuan Zhang, Jane H. Lee, Shivani AgarwalICML 2021 · 38 citations
- QuASE: Question-Answer Driven Sentence EncodingHangfeng He, Qiang Ning, Dan RothACL 2020 · 32 citations
- Learnability with Indirect Supervision SignalsKaifu Wang, Qiang Ning, Dan RothNeurIPS 2020 · 10 citations
Related papers
- Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State AccessDaniel Ebi, Damien Ernst, Klemens Böhm, Gaspard LambrechtsICML 2026 · 1 citation
- Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-AnsweringYavuz Faruk Bakman, Sungmin Kang, Zhiqi Huang, Duygu Nur Yaldiz et al.ICLR 2026 · 6 citations
- Towards Explainable Joint Models via Information Theory for Multiple Intent Detection and Slot FillingXianwei Zhuang, Xuxin Cheng, Yuexian ZouAAAI 2024 · 23 citations
- Unifying Knowledge Base Completion with PU Learning to Mitigate the Observation BiasJonas Schouterden, Jessa Bekker, Jesse Davis, Hendrik BlockeelAAAI 2022 · 2 citations
- A Bayesian Framework for Information-Theoretic ProbingTiago Pimentel, Ryan CotterellEMNLP 2021 · 2 citations
