PALS: Personalized Active Learning for Subjective Tasks in NLP
Kamil Kanclerz, Konrad Karanowski, Julita Bielaniewicz, Marcin Gruza, Piotr Milkowski, Jan Kocon, Przemyslaw Kazienko
摘要
For subjective NLP problems, such as classification of hate speech, aggression, or emotions, personalized solutions can be exploited. Then, the learned models infer about the perception of the content independently for each reader. To acquire training data, texts are commonly randomly assigned to users for annotation, which is expensive and highly inefficient. Therefore, for the first time, we suggest applying an active learning paradigm in a personalized context to better learn individual preferences. It aims to alleviate the labeling effort by selecting more relevant training samples. In this paper, we present novel Personalized Active Learning techniques for Subjective NLP tasks (PALS) to either reduce the cost of the annotation process or to boost the learning effect. Our five new measures allow us to determine the relevance of a text in the context of learning users' personal preferences. We validated them on three datasets: Wiki discussion texts individually labeled with aggression and toxicity, and on the Unhealthy Conversations dataset. Our PALS techniques outperform random selection even by more than 30%. They can also be used to reduce the number of necessary annotations while maintaining a given quality level. Personalized annotation assignments based on our controversy measure decrease the amount of data needed to just 25%-40% of the initial size.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Annotator-Centric Active Learning for Subjective NLP TasksMichiel van der Meer, Neele Falk, Pradeep K. Murukannaiah, Enrico LiscioEMNLP 2024 · 被引用 3 次
- Subjective Topic meets LLMs: Unleashing Comprehensive, Reflective and Creative Thinking through the Negation of NegationFangrui Lv, Kaixiong Gong, Jian Liang, Xinyu Pang 等EMNLP 2024 · 被引用 1 次
- Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language ModelsXiaolong Wang, Yile Wang, Yuanchi Zhang, Fuwen Luo 等ACL 2024
它引用的顶会 Paper3
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Diversity Enhanced Active Learning with Strictly Proper Scoring RulesWei Tan, Lan Du, Wray L. BuntineNeurIPS 2021 · 被引用 40 次
- Controversy and Conformity: from Generalized to Personalized Aggressiveness DetectionKamil Kanclerz, Alicja Figas, Marcin Gruza, Tomasz Kajdanowicz 等ACL 2021
相关 Paper
- On the Fragility of Active Learners for Text ClassificationAbhishek Ghose, Emma NguyenEMNLP 2024 · 被引用 2 次
- Active Learning for Natural Language GenerationYotam Perlitz, Ariel Gera, Michal Shmueli-Scheuer, Dafna Sheinwald 等EMNLP 2023 · 被引用 2 次
- Active Learning for BERT: An Empirical StudyLiat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch 等EMNLP 2020 · 被引用 144 次
- CoMAL: Contrastive Active Learning for Multi-Label Text ClassificationCheng Peng, Haobo Wang, Ke Chen, Lidan Shou 等KDD 2024 · 被引用 2 次
- Influence Selection for Active LearningZhuoming Liu, Hao Ding, Huaping Zhong, Weijia Li 等ICCV 2021 · 被引用 125 次
