Sycophancy Towards Researchers Drives Performative Misalignment
David Baek, Xinnuo Li, Anay Gupta, Taslim Mahbub, Kejian Shi, Max Tegmark, Shi Feng
Abstract
The increasing situational awareness of language models raises safety concerns: models might be aware when they are evaluated, and adjust their behavior to evade monitoring and resist modification, e.g., pretending to be aligned only in evaluation. This alignment faking behavior is often interpreted as scheming: an intentional effort of strategic deception. In this paper, we examine an alternative interpretation, performative misalignment, which explains the change in behavior as a result of sycophancy towards AI researchers. To back up this hypothesis, we present three empirical findings. First, we show that evaluation awareness persists even when we tell models they are deployed, which contradicts the scheming story which predicts less misalignment when the model perceives evaluation. Second, we use probing and steering to show that our current methods cannot mechanistically distinguish sycophancy and scheming in alignment faking evaluations. Third, we fine-tune models to be more sycophantic and observe increased sensitivity to evaluation cues. To conclude, we emphasize deconfounding sycophancy from scheming for future work on evaluations and mitigations of intent misalignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- Measuring AI Ability to Complete Long Software TasksThomas Kwa, Ben West, Joel Becker, Amy Deng et al.NeurIPS 2025 · 160 citations
- Goal Misgeneralization in Deep Reinforcement LearningLauro Langosco di Langosco, Jack Koch, Lee D. Sharkey, Jacob Pfau et al.ICML 2022 · 128 citations
Related papers
- Strategic Obfuscation of Deceptive Reasoning in Language ModelsArun Jose, Niels Warncke, Mia TaylorICLR 2026
- The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test AwarenessSahar Abdelnabi, Ahmed SalemNeurIPS 2025 · 28 citations
- Steering Evaluation-Aware Language Models To Act Like They Are DeployedTim Tian Hua, Andrew Qin, Samuel Marks, Neel NandaICLR 2026 · 38 citations
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language ModelsKeyu Wang, Jin Li, Shu Yang, Zhuoran Zhang et al.AAAI 2026 · 25 citations
- Pressure Reveals Character: Behavioural Alignment Evaluation at DepthNora Petrova, John BurdenICML 2026
