Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis
Guangliang Liu, Haitao Mao, Jiliang Tang, Kristen Marie Johnson
Abstract
Large Language Models (LLMs) are capable of producing content that perpetuates stereotypes, discrimination, and toxicity. The recently proposed moral self-correction is a computationally efficient method for reducing harmful content in the responses of LLMs. However, the process of how injecting self-correction instructions can modify the behavior of LLMs remains under-explored. In this paper, we explore the effectiveness of moral self-correction by answering three research questions: ( 1 ) In what scenarios does moral self-correction work? (2) What are the internal mechanisms of LLMs, e.g., hidden states, that are influenced by moral selfcorrection instructions? (3) Is intrinsic moral self-correction actually superficial in terms of reduced immorality in hidden states? We argue that self-correction can help LLMs find a shortcut to more morally correct output, rather than truly reducing the immorality stored in hidden states. Through empirical investigation with tasks of language generation and multi-choice question answering, we conclude: (i) LLMs exhibit good performance across both tasks, and self-correction instructions are particularly beneficial when the correct answer is already top-ranked; (ii) The morality levels in intermediate hidden states are strong indicators as to whether one instruction would be more effective than another; (iii) Based on our analysis of intermediate hidden states and task case studies of self-correction behaviors, we are first to propose the hypothesis that intrinsic moral self-correction is in fact superficial.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a097534b-f44b-459e-a8c2-07932db6460aCited by top-tier papers3
- Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMsAngelina Wang, Michelle Phan, Daniel E. Ho, Sanmi KoyejoACL 2025 · 17 citations
- The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate ItLeonardo Bertolazzi, Philipp Mondorf, Barbara Plank, Raffaella BernardiEMNLP 2025 · 8 citations
- Do Morals Guide How LLMs Think? The Role of Ethical Perspectives in General Problem SolvingIseo Kim, Eunjin Hong, Juae KimACL 2026
Builds on16
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
Related papers
- Small Language Model Can Self-CorrectHaixia Han, Jiaqing Liang, Jie Shi, Qianyu He et al.AAAI 2024 · 31 citations
- Understanding the Dark Side of LLMs' Intrinsic Self-CorrectionQingjie Zhang, Di Wang, Haoting Qian, Yiming Li et al.ACL 2025 · 36 citations
- On Large Language Models' Resilience to Coercive InterrogationZhuo Zhang, Guangyu Shen, Guanhong Tao, Siyuan Cheng et al.S&P 2024 · 24 citations
- SelfIE: Self-Interpretation of Large Language Model EmbeddingsHaozhe Chen, Carl Vondrick, Chengzhi MaoICML 2024 · 58 citations
- Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender DiscourseRongchen Guo, Isar Nejadgholi, Hillary Dawkins, Kathleen C. Fraser et al.EMNLP 2024
