Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data
Faeze Ghorbanpour, Daryna Dementieva, Alexander Fraser
摘要
Considering the importance of detecting hateful content, labeled hate speech data is expensive and time-consuming to collect and annotate, particularly for low-resource languages.Prior work has demonstrated the effectiveness of cross-lingual transfer learning and data augmentation in improving performance on tasks with limited labeled data.To develop an efficient and scalable cross-lingual transfer learning approach, we leverage nearest-neighbor retrieval to augment minimal labeled data in the target language, thereby enhancing detection performance.Specifically, we assume access to a small set of labeled training instances in the target language and use these to retrieve the most relevant labeled examples from a large multilingual hate speech detection pool.We evaluate our approach on eight languages and demonstrate that it consistently outperforms models trained solely on the target language data.Furthermore, in most cases, our method surpasses the current state-of-the-art.Notably, our approach is highly data-efficient, retrieving as few as 200 instances in some cases while maintaining superior performance.Moreover, it is scalable, as the retrieval pool can be easily expanded, and the method can be readily adapted to new languages and tasks.We also apply maximum marginal relevance to mitigate redundancy and filter out highly similar retrieved instances, resulting in improvements in some languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi 等EMNLP 2021 · 被引用 159 次
相关 Paper
- Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced LanguagesPaul Röttger, Debora Nozza, Federico Bianchi, Dirk HovyEMNLP 2022 · 被引用 16 次
- Vicinal Risk Minimization for Few-Shot Cross-lingual Transfer in Abusive Language DetectionGretel Liz De la Peña Sarracén, Paolo Rosso, Robert Litschko, Goran Glavas 等EMNLP 2023 · 被引用 2 次
- kNN-TL: k-Nearest-Neighbor Transfer Learning for Low-Resource Neural Machine TranslationShudong Liu, Xuebo Liu, Derek F. Wong, Zhaocong Li 等ACL 2023 · 被引用 14 次
- Multilingual Meta-Distillation Alignment for Semantic RetrievalMeryem M'hamdi, Jonathan May, Franck Dernoncourt, Trung Bui 等SIGIR 2024
- UXLA: A Robust Unsupervised Data Augmentation Framework for Zero-Resource Cross-Lingual NLPM. Saiful Bari, Tasnim Mohiuddin, Shafiq R. JotyACL 2021
