Explaining Matters: Leveraging Definitions and Semantic Expansion for Sexism Detection
Sahrish Khan, Arshad Jhumka, Gabriele Pergola
摘要
The detection of sexism in online content remains an open problem, as harmful language disproportionately affects women and marginalized groups. While automated systems for sexism detection have been developed, they still face two key challenges: data sparsity and the nuanced nature of sexist language. Even in large, well-curated datasets like the Explainable Detection of Online Sexism (EDOS), severe class imbalance hinders model generalization. Additionally, the overlapping and ambiguous boundaries of fine-grained categories introduce substantial annotator disagreement, reflecting the difficulty of interpreting nuanced expressions of sexism. To address these challenges, we propose two prompt-based data augmentation techniques: Definition-based Data Augmentation (DDA), which leverages categoryspecific definitions to generate semanticallyaligned synthetic examples, and Contextual Semantic Expansion (CSE), which targets systematic model errors by enriching examples with task-specific semantic features. To further improve reliability in fine-grained classification, we introduce an ensemble strategy that resolves prediction ties by aggregating complementary perspectives from multiple language models. Our experimental evaluation on the EDOS dataset demonstrates state-of-the-art performance across all tasks, with notable improvements of macro F1 by 1.5 points for binary classification (Task A) and 4.1 points for fine-grained classification (Task C) 1 . Warning: This paper includes examples that might be offensive and upsetting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Verify-and-Edit: A Knowledge-Enhanced Chain-of-Thought FrameworkRuochen Zhao, Xingxuan Li, Shafiq Joty, Chengwei Qin 等ACL 2023 · 被引用 50 次
- Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model CommunicationZhangyue Yin, Qiushi Sun, Cheng Chang, Qipeng Guo 等EMNLP 2023 · 被引用 15 次
- XtremeDistil: Multi-stage Distillation for Massive Multilingual ModelsSubhabrata Mukherjee, Ahmed Hassan AwadallahACL 2020 · 被引用 4 次
- Voices in a Crowd: Searching for clusters of unique perspectivesNikolas Vitsakis, Amit Parekh, Ioannis KonstasEMNLP 2024
相关 Paper
- BeyondGender: A Multifaceted Bilingual Dataset for Practical Sexism DetectionXuan Luo, Li Yang, Han Zhang, Geng Tu 等AAAI 2025 · 被引用 3 次
- RoPDA: Robust Prompt-Based Data Augmentation for Low-Resource Named Entity RecognitionSihan Song, Furao Shen, Jian ZhaoAAAI 2024 · 被引用 7 次
- Annotating Online MisogynyPhiline Zeinert, Nanna Inie, Leon DerczynskiACL 2021
- Prompt Tuning Pushes Farther, Contrastive Learning Pulls Closer: A Two-Stage Approach to Mitigate Social BiasesYingji Li, Mengnan Du, Xin Wang, Ying WangACL 2023 · 被引用 12 次
- PromDA: Prompt-based Data Augmentation for Low-Resource NLU TasksYufei Wang, Can Xu, Qingfeng Sun, Huang Hu 等ACL 2022
