When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
Keyu Wang, Jin Li, Shu Yang, Zhuoran Zhang, Di Wang
Abstract
Large Language Models (LLMs) often exhibit sycophantic behavior, agreeing with user-stated opinions even when those contradict factual knowledge. While prior work has documented this tendency, the internal mechanisms that enable such behavior remain poorly understood. In this paper, we provide a mechanistic account of how sycophancy arises within LLMs. We first systematically study how user opinions induce sycophancy across different model families. We find that simple opinion statements reliably induce sycophancy, whereas user expertise framing has a negligible impact. Through logit-lens analysis and causal activation patching, we identify a two-stage emergence of sycophancy: (1) a late-layer output preference shift and (2) deeper representational divergence. We also verify that user authority fails to influence behavior, because models do not encode it internally. In addition, we examine how grammatical perspective affects sycophantic behavior, finding that first-person prompts ("I believe...") consistently induce higher sycophancy rates than third-person framings ("They believe...") by creating stronger representational perturbations in deeper layers. These findings highlight that sycophancy is not a surface-level artifact but emerges from a structural override of learned knowledge in deeper layers, with implications for alignment and truthful AI systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da2b69b8-0926-4e07-b97a-da3f471e05beCited by top-tier papers16
- EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit IdentificationLin Zhang, Wenshuo Dong, Zhuoran Zhang, Shu Yang et al.NeurIPS 2025 · 26 citations
- Interaction Context Often Increases Sycophancy in LLMsShomik Jain, Charlotte Park, Matt Viana, Ashia Wilson et al.CHI 2026 · 12 citations
- When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent ReasoningHyeong Kyu Choi, Xiaojin (Jerry) Zhu, Sharon LiACL 2026 · 12 citations
- PIXEL: Adaptive Steering Via Position-wise Injection with eXact Estimated Levels under a Subspace CalibrationManjiang Yu, Hongji Li, Priyanka Singh, Xue Li et al.WWW 2026 · 11 citations
- Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving TasksJessica Y. Bo, Majeed Kazemitabaar, Mengqing Deng, Michael Inzlicht et al.CHI 2026 · 8 citations
Builds on19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
Related papers
- Causally Motivated Sycophancy Mitigation for Large Language ModelsHaoxi Li, Xueyang Tang, Jie Zhang, Song Guo et al.ICLR 2025
- How RLHF Amplifies SycophancyItai Shapira, Gerdus Benade, Ariel ProcacciaICML 2026 · 16 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language ModelsArya Shah, Deepali Mishra, Chaklam SilpasuwanchaiACL 2026 · 1 citation
- Have the VLMs Lost Confidence? A Study of Sycophancy in VLMsShuo Li, Tao Ji, Xiaoran Fan, Linsheng Lu et al.ICLR 2025
