Vicinal Risk Minimization for Few-Shot Cross-lingual Transfer in Abusive Language Detection
Gretel Liz De la Peña Sarracén, Paolo Rosso, Robert Litschko, Goran Glavas, Simone Paolo Ponzetto
摘要
Cross-lingual transfer learning from high-resource to medium and low-resource languages has shown encouraging results. However, the scarcity of resources in target languages remains a challenge. In this work, we resort to data augmentation and continual pre-training for domain adaptation to improve cross-lingual abusive language detection. For data augmentation, we analyze two existing techniques based on vicinal risk minimization and propose MIXAG, a novel data augmentation method which interpolates pairs of instances based on the angle of their representations. Our experiments involve seven languages typologically distinct from English and three different domains. The results reveal that the data augmentation strategies can enhance few-shot cross-lingual abusive language detection. Specifically, we observe that consistently in all target languages, MIXAG improves significantly in multidomain and multilingual environments. Finally, we show through an error analysis how the domain adaptation can favour the class of abusive texts (reducing false negatives), but at the same time, declines the precision of the abusive language detection model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts 等ACL 2023 · 被引用 319 次
- FlipDA: Effective and Robust Data Augmentation for Few-Shot LearningJing Zhou, Yanan Zheng, Jie Tang, Li Jian 等ACL 2022 · 被引用 91 次
- Don't Stop Fine-Tuning: On Training Regimes for Few-Shot Cross-Lingual Transfer with Multilingual Language ModelsFabian David Schmidt, Ivan Vulic, Goran GlavasEMNLP 2022 · 被引用 12 次
- SSMBA: Self-Supervised Manifold Based Data Augmentation for Improving Out-of-Domain RobustnessNathan Ng, Kyunghyun Cho, Marzyeh GhassemiEMNLP 2020 · 被引用 6 次
相关 Paper
- Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled DataFaeze Ghorbanpour, Daryna Dementieva, Alexander FraserEMNLP 2025
- UXLA: A Robust Unsupervised Data Augmentation Framework for Zero-Resource Cross-Lingual NLPM. Saiful Bari, Tasnim Mohiuddin, Shafiq R. JotyACL 2021
- Data Augmentation with Adversarial Training for Cross-Lingual NLIXin Dong, Yaxin Zhu, Zuohui Fu, Dongkuan Xu 等ACL 2021
- Multilingual Transfer Learning for QA using Translation as Data AugmentationMihaela A. Bornea, Lin Pan, Sara Rosenthal, Radu Florian 等AAAI 2021 · 被引用 45 次
- Culture Matters in Toxic Language Detection in PersianZahra Bokaei, Walid Magdy, Bonnie WebberACL 2025
