Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse
Rongchen Guo, Isar Nejadgholi, Hillary Dawkins, Kathleen C. Fraser, Svetlana Kiritchenko
Abstract
This work provides an explanatory view of how LLMs can apply moral reasoning to both criticize and defend sexist language. We assessed eight large language models, all of which demonstrated the capability to provide explanations grounded in varying moral perspectives for both critiquing and endorsing views that reflect sexist assumptions. With both human and automatic evaluation, we show that all eight models produce comprehensible and contextually relevant text, which is helpful in understanding diverse views on how sexism is perceived. Also, through analysis of moral foundations cited by LLMs in their arguments, we uncover the diverse ideological perspectives in models' outputs, with some models aligning more with progressive or conservative views on gender roles and sexism. Based on our observations, we caution against the potential misuse of LLMs to justify sexist language. We also highlight that LLMs can serve as tools for understanding the roots of sexist beliefs and designing well-informed interventions. Given this dual capacity, it is crucial to monitor LLMs and design safety mechanisms for their use in applications that involve sensitive societal topics, such as sexism. Warning: This paper includes examples that might be offensive and upsetting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f10ed2d0-da65-421c-b475-393cef63d130Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- Differentiable Prompt Makes Pre-trained Language Models Better Few-shot LearnersNingyu Zhang, Luoqiu Li, Xiang Chen, Shumin Deng et al.ICLR 2022 · 205 citations
- Counterspeakers' Perspectives: Unveiling Barriers and AI Needs in the Fight against Online HateJimin Mun, Cathy Buerger, Jenny T. Liang, Joshua Garland et al.CHI 2024 · 12 citations
- Unintended Impacts of LLM Alignment on Global RepresentationMichael J. Ryan, William Barr Held, Diyi YangACL 2024
Related papers
- A Matter of Perspective(s): Contrasting Human and LLM Argumentation in Subjective Decision-Making on Subtle SexismPaula Akemi Aoyagui, Kelsey Stemmler, Sharon A. Ferguson, Young-Ho Kim et al.CHI 2025 · 4 citations
- Language is Scary when Over-Analyzed: Unpacking Implied Misogynistic Reasoning with Argumentation Theory-Driven PromptsArianna Muti, Federico Ruggeri, Khalid Al-Khatib, Alberto Barrón-Cedeño et al.EMNLP 2024 · 1 citation
- SOLAR: Towards Characterizing Subjectivity of Individuals through Modeling Value Conflicts and Trade-offsYounghun Lee, Dan GoldwasserEMNLP 2025
- Do Morals Guide How LLMs Think? The Role of Ethical Perspectives in General Problem SolvingIseo Kim, Eunjin Hong, Juae KimACL 2026
- Moral Foundations of Large Language ModelsMarwa Abdulhai, Gregory Serapio-García, Clément Crepy, Daria Valter et al.EMNLP 2024 · 22 citations
