Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data
Faeze Ghorbanpour, Daryna Dementieva, Alexander Fraser
Abstract
Considering the importance of detecting hateful content, labeled hate speech data is expensive and time-consuming to collect and annotate, particularly for low-resource languages.Prior work has demonstrated the effectiveness of cross-lingual transfer learning and data augmentation in improving performance on tasks with limited labeled data.To develop an efficient and scalable cross-lingual transfer learning approach, we leverage nearest-neighbor retrieval to augment minimal labeled data in the target language, thereby enhancing detection performance.Specifically, we assume access to a small set of labeled training instances in the target language and use these to retrieve the most relevant labeled examples from a large multilingual hate speech detection pool.We evaluate our approach on eight languages and demonstrate that it consistently outperforms models trained solely on the target language data.Furthermore, in most cases, our method surpasses the current state-of-the-art.Notably, our approach is highly data-efficient, retrieving as few as 200 instances in some cases while maintaining superior performance.Moreover, it is scalable, as the retrieval pool can be easily expanded, and the method can be readily adapted to new languages and tasks.We also apply maximum marginal relevance to mitigate redundancy and filter out highly similar retrieved instances, resulting in improvements in some languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi et al.EMNLP 2021 · 159 citations
Related papers
- Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced LanguagesPaul Röttger, Debora Nozza, Federico Bianchi, Dirk HovyEMNLP 2022 · 16 citations
- Vicinal Risk Minimization for Few-Shot Cross-lingual Transfer in Abusive Language DetectionGretel Liz De la Peña Sarracén, Paolo Rosso, Robert Litschko, Goran Glavas et al.EMNLP 2023 · 2 citations
- kNN-TL: k-Nearest-Neighbor Transfer Learning for Low-Resource Neural Machine TranslationShudong Liu, Xuebo Liu, Derek F. Wong, Zhaocong Li et al.ACL 2023 · 14 citations
- Multilingual Meta-Distillation Alignment for Semantic RetrievalMeryem M'hamdi, Jonathan May, Franck Dernoncourt, Trung Bui et al.SIGIR 2024
- UXLA: A Robust Unsupervised Data Augmentation Framework for Zero-Resource Cross-Lingual NLPM. Saiful Bari, Tasnim Mohiuddin, Shafiq R. JotyACL 2021
