BeyondGender: A Multifaceted Bilingual Dataset for Practical Sexism Detection
Xuan Luo, Li Yang, Han Zhang, Geng Tu, Qianlong Wang, Keyang Ding, Chuang Fan, Jing Li, Ruifeng Xu
Abstract
Sexism affects both women and men, yet research often overlooks misandry and suffers from overly broad annotations that limit AI applications. To address this, we introduce Be-yondGender, a dataset meticulously annotated according to the latest definitions of misogyny and misandry. It features innovative multifaceted labels encompassing aspects of sexism, gender, phrasing, misogyny, and misandry. The dataset includes 6.0K English and 1.7K Chinese sexism instances, alongside 13.4K non-sexism examples. Our evaluations of masked language models and large language models reveal that they detect misogyny in English and misandry in Chinese more effectively, with F1-scores of 0.87 and 0.62, respectively. However, they frequently misclassify hostile and mild comments, underscoring the complexity of sexism detection. Parallel corpus experiments suggest promising data augmentation strategies to enhance AI systems for nuanced sexism detection, and our dataset can be leveraged to improve value alignment in large language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa0042ee-fa8e-4c41-9d03-359b128e05beBuilds on3
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- Hate-Speech and Offensive Language Detection in Roman UrduHammad Rizwan, Muhammad Haroon Shakeel, Asim KarimEMNLP 2020 · 97 citations
- Annotating Online MisogynyPhiline Zeinert, Nanna Inie, Leon DerczynskiACL 2021
Related papers
- GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language ModelsTao Zhang, Ziqian Zeng, Yuxiang Xiao, Huiping Zhuang et al.ACL 2025 · 18 citations
- Multi-Dimensional Gender Bias ClassificationEmily Dinan, Angela Fan, Ledell Wu, Jason Weston et al.EMNLP 2020 · 7 citations
- FairPrism: Evaluating Fairness-Related Harms in Text GenerationEve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett et al.ACL 2023 · 9 citations
- Gender Bias in Multilingual Embeddings and Cross-Lingual TransferJieyu Zhao, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang et al.ACL 2020 · 59 citations
- Language is Scary when Over-Analyzed: Unpacking Implied Misogynistic Reasoning with Argumentation Theory-Driven PromptsArianna Muti, Federico Ruggeri, Khalid Al-Khatib, Alberto Barrón-Cedeño et al.EMNLP 2024 · 1 citation
