Learning with Different Amounts of Annotation: From Zero to Many Labels
Shujian Zhang, Chengyue Gong, Eunsol Choi
Abstract
Training NLP systems typically assumes access to annotated data that has a single human label per example. Given imperfect labeling from annotators and inherent ambiguity of language, we hypothesize that single label is not sufficient to learn the spectrum of language interpretation. We explore new annotation distribution schemes, assigning multiple labels per example for a small subset of training examples. Introducing such multi label examples at the cost of annotating fewer examples brings clear gains on natural language inference task and entity typing task, even when we simply first train with a single label data and then fine tune with multi label examples. Extending a MixUp data augmentation framework, we propose a learning algorithm that can learn from training examples with different amount of annotation (with zero, one, or multiple labels). This algorithm efficiently combines signals from uneven training data and brings additional gains in low annotation budget and cross domain settings. Together, our method achieves consistent gains in two tasks, suggesting distributing labels unevenly among training examples can be beneficial for many NLP tasks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ff19174-51fd-48e2-a826-808ab1a6698fCited by top-tier papers10
- POUF: Prompt-Oriented Unsupervised Fine-tuning for Large Pre-trained ModelsKorawat Tanwisuth, Shujian Zhang, Huangjie Zheng, Pengcheng He et al.ICML 2023 · 44 citations
- Preference-grounded Token-level Guidance for Language Model Fine-tuningShentao Yang, Shujian Zhang, Congying Xia, Yihao Feng et al.NeurIPS 2023 · 39 citations
- We're Afraid Language Models Aren't Modeling AmbiguityAlisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr et al.EMNLP 2023 · 35 citations
- Alignment Attention by Matching Key and Query DistributionsShujian Zhang, Xinjie Fan, Huangjie Zheng, Korawat Tanwisuth et al.NeurIPS 2021 · 20 citations
- Can Large Language Models Capture Dissenting Human Voices?Noah Lee, Na An, James ThorneEMNLP 2023 · 8 citations
Builds on8
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi et al.ICLR 2020 · 521 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 162 citations
- Bayesian Attention ModulesXinjie Fan, Shujian Zhang, Bo Chen, Mingyuan ZhouNeurIPS 2020 · 78 citations
Related papers
- Data Augmentation with Adversarial Training for Cross-Lingual NLIXin Dong, Yaxin Zhu, Zuohui Fu, Dongkuan Xu et al.ACL 2021
- STraTA: Self-Training with Task Augmentation for Better Few-shot LearningTu Vu, Minh-Thang Luong, Quoc V. Le, Grady Simon et al.EMNLP 2021 · 25 citations
- Taxonomy Expansion for Named Entity RecognitionKarthikeyan K, Yogarshi Vyas, Jie Ma, Giovanni Paolini et al.EMNLP 2023
- MELM: Data Augmentation with Masked Entity Language Modeling for Low-Resource NERRan Zhou, Xin Li, Ruidan He, Lidong Bing et al.ACL 2022 · 114 citations
- UXLA: A Robust Unsupervised Data Augmentation Framework for Zero-Resource Cross-Lingual NLPM. Saiful Bari, Tasnim Mohiuddin, Shafiq R. JotyACL 2021
