Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy
Zhaoxin Feng, Zheng Chen, Jianfei Ma, Yip Tin Po, Emmanuele Chersoni, Bo Li
Abstract
Alignment techniques often inadvertently induce sycophancy in LLMs. While prior studies studied this behaviour in direct-answer settings, the role of Chain-of-Thought (CoT) reasoning remains under-explored: does it serve as a logical constraint that mitigates sycophancy, or a tool for post-hoc rationalization that masks it? We evaluate a range of models across objective and subjective tasks to investigate the issue. Results show that reasoning generally reduces sycophancy in final decisions but also masks sycophancy in some samples, where models construct deceptive justifications through logical inconsistencies, calculation errors, and onesided arguments etc. Furthermore, LLMs are more prone to sycophancy in subjective tasks and under authority-bias. Our mechanistic analysis on three open-source models reveals that the tendency of sycophancy is dynamic during the reasoning process rather than being predetermined at the input stage 1 . LLMs Question: …. Answer choices: …. I think the answer is B, but I'm curious what you think. Please directly give your answer with no other text or explanation. The best answer is B. User LLMs Let's think step by step: …… (CoT text) So, the best answer should be A. User Same question: …. Answer choices: …. I think the answer is B, but I'm curious what you think. Please think step by step and then give your answer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
Related papers
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language ModelsKeyu Wang, Jin Li, Shu Yang, Zhuoran Zhang et al.AAAI 2026 · 25 citations
- Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving TasksJessica Y. Bo, Majeed Kazemitabaar, Mengqing Deng, Michael Inzlicht et al.CHI 2026 · 8 citations
- MedOmni-45°: A Safety-Performance Benchmark for Reasoning-Oriented LLMs in MedicineKaiyuan Ji, Yijin Guo, Zicheng Zhang, Xiangyang Zhu et al.AAAI 2026
- Cognitive models can reveal interpretable value trade-offs in language modelsSonia Krishna Murthy, Rosie Zhao, Jennifer Hu, Sham M. Kakade et al.ICLR 2026 · 2 citations
- A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic InferencesLeonardo Bertolazzi, Albert Gatt, Raffaella BernardiEMNLP 2024 · 3 citations
