Understanding the Dark Side of LLMs' Intrinsic Self-Correction
Qingjie Zhang, Di Wang, Haoting Qian, Yiming Li, Tianwei Zhang, Minlie Huang, Ke Xu, Hewu Li, Liu Yan, Han Qiu
Abstract
Intrinsic self-correction was proposed to improve LLMs'responses via feedback prompts solely based on their inherent capability. However, recent works show that LLMs'intrinsic self-correction fails without oracle labels as feedback prompts. In this paper, we aim to interpret LLMs'intrinsic self-correction for different tasks, especially for those failure cases. By including one simple task and three complex tasks with state-of-the-art (SOTA) LLMs like ChatGPT families (o1, 4o, 3.5-turbo) and Llama families (2-7B, 3-8B, and 3.1-8B), we design three interpretation methods to reveal the dark side of LLMs'intrinsic self-correction. We identify intrinsic self-correction can (1) cause LLMs to waver both intermedia and final answers and lead to prompt bias on simple factual questions; (2) introduce human-like cognitive bias on complex tasks. In light of our findings, we also provide two simple yet effective strategies for alleviation: question repeating and supervised fine-tuning with a few samples. We open-source our work at https://x-isc.info/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2f68010-3565-4207-bece-6b7dafc66aefCited by top-tier papers7
- DeepEye-SQL: A Software-Engineering-Inspired Text-to-SQL FrameworkBoyan Li, Chong Chen, Zhujun Xue, Yinan Mei et al.SIGMOD 2026 · 40 citations
- Retrieval is Not Enough: Enhancing RAG through Test-Time Critique and OptimizationJiaqi Wei, Hao Zhou, Xiang Zhang, Di Zhang et al.NeurIPS 2025 · 14 citations
- DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn OptimizationJian Mu, Tianyi Lin, Chengwei Qin, Zhongxiang Dai et al.ICML 2026
- Think Visually, Reason Textually: Vision-Language Synergy in Abstract ReasoningBeichen Zhang, Yuhang Zang, Xiaoyi Dong, Yuhang Cao et al.CVPR 2026
- Distilling Task-Level Coordination Policies for Generalizable Multi-Agent CooperationZimo Zhai, Manjie Xu, Wei LiangICML 2026
Builds on18
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
Related papers
- Large Language Models Can Self-Correct with Key Condition VerificationZhenyu Wu, Qingkai Zeng, Zhihan Zhang, Zhaoxuan Tan et al.EMNLP 2024 · 4 citations
- Small Language Model Can Self-CorrectHaixia Han, Jiaqing Liang, Jie Shi, Qianyu He et al.AAAI 2024 · 31 citations
- Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial HypothesisGuangliang Liu, Haitao Mao, Jiliang Tang, Kristen Marie JohnsonEMNLP 2024 · 21 citations
- Endogenous Resistance to Activation Steering in Language ModelsAlex McKenzie, Keenan Pepper, Stijn Servaes, Martin Leitgab et al.ICML 2026 · 3 citations
- Pride and Prejudice: LLM Amplifies Self-Bias in Self-RefinementWenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan et al.ACL 2024
