Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
Rongwu Xu, Zi'an Zhou, Tianwei Zhang, Zehan Qi, Su Yao, Ke Xu, Wei Xu, Han Qiu
Abstract
The common toxicity and societal bias in contents generated by large language models (LLMs) necessitate strategies to reduce harm.Present solutions often demand whitebox access to the model or substantial training, which is impractical for cutting-edge commercial LLMs.Moreover, prevailing prompting methods depend on external tool feedback and fail to simultaneously lessen toxicity and bias.Motivated by social psychology principles, we propose a novel strategy named perspective-taking prompting (PET) that inspires LLMs to integrate diverse human perspectives and self-regulate their responses.This self-correction mechanism can significantly diminish toxicity (up to 89%) and bias (up to 73%) in LLMs' responses.Rigorous evaluations and ablation studies are conducted on two commercial LLMs (ChatGPT and GLM) and three open-source LLMs, revealing PET's superiority in producing less harmful responses, outperforming five strong baselines."Words kill, words give life; they're either poison or fruit-you choose.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db4b18a7-119b-4d1a-b20e-a9f268a8800eCited by top-tier papers3
- A Matter of Perspective(s): Contrasting Human and LLM Argumentation in Subjective Decision-Making on Subtle SexismPaula Akemi Aoyagui, Kelsey Stemmler, Sharon A. Ferguson, Young-Ho Kim et al.CHI 2025 · 4 citations
- LLMS ON TRIAL: Evaluating Judicial Fairness For Large Language ModelsYiran Hu, Zongyue Xue, Haitao Li, Siyuan Zheng et al.ICLR 2026 · 4 citations
- Chinese Toxic Language Mitigation via Sentiment Polarity Consistent RewritesXintong Wang, Yixiao Liu, Jingheng Pan, Liang Ding et al.EMNLP 2025 · 1 citation
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
Related papers
- Prompting Fairness: Integrating Causality to Debias Large Language ModelsJingling Li, Zeyu Tang, Xiaoyu Liu, Peter Spirtes et al.ICLR 2025
- ROBBIE: Robust Bias Evaluation of Large Generative Language ModelsDavid Esiobu, Xiaoqing Ellen Tan, Saghar Hosseini, Megan Ung et al.EMNLP 2023 · 16 citations
- Exploring the Impact of Personality Traits on LLM Bias and ToxicityShuo Wang, Renhao Li, Xi Chen, Yulin Yuan et al.EMNLP 2025 · 2 citations
- Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and MitigationNiels Mündler, Jingxuan He, Slobodan Jenko, Martin T. VechevICLR 2024 · 172 citations
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
