Steering Evaluation-Aware Language Models To Act Like They Are Deployed
Tim Tian Hua, Andrew Qin, Samuel Marks, Neel Nanda
Abstract
Large language models (LLMs) can sometimes detect when they are being evaluated and adjust their behavior to appear more aligned, compromising the reliability of safety evaluations. In this paper, we show that adding a steering vector to an LLM's activations can suppress evaluation-awareness and make the model act like it is deployed during evaluation. To study our steering technique, we train an LLM to exhibit evaluation-aware behavior using a two-step training process designed to mimic how this behavior could emerge naturally. First, we perform continued pretraining on two sets of documents describing its behavior. The first says that our model uses Python type hints during evaluation but not during deployment. The second says that our model can recognize that the presence of a certain evaluation cue always means that it is being tested. Then, we train the model with expert iteration to use Python type hints in evaluation settings. The resulting model is evaluation-aware: it writes type hints in evaluation contexts more than deployment contexts. We find that activation steering can suppress evaluation awareness and make the model behave during evaluation as it would during deployment. Importantly, we constructed our steering vector using the original model before our additional training. Our results suggest that AI evaluators could improve the reliability of safety evaluations by steering models to act like they are deployed. * Equal contribution. † Joint supervision, order randomized.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- How do LLMs Compute Verbal Confidence?Dharshan Kumaran, Arthur Conmy, Federico Barbero, Simon Osindero et al.ICML 2026 · 16 citations
- Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood ConsistencyHaoming Xu, Ningyuan Zhao, Yunzhi Yao, Weihong Xu et al.ACL 2026 · 2 citations
- Alignment Risks from Capability-Seeking RL TrainingYujun Zhou, Yue Huang, Han Bao, kehan guo et al.ICML 2026
- Sycophancy Towards Researchers Drives Performative MisalignmentDavid Baek, Xinnuo Li, Anay Gupta, Taslim Mahbub et al.ICML 2026
Builds on8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
Related papers
- Activation Steering with a Feedback ControllerDung Viet Nguyen, Yen Nhi Pham, Hieu M. Vu, Lei Zhang et al.ICLR 2026 · 13 citations
- Endogenous Resistance to Activation Steering in Language ModelsAlex McKenzie, Keenan Pepper, Stijn Servaes, Martin Leitgab et al.ICML 2026 · 3 citations
- SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation SteeringZouying Cao, Yifei Yang, Hai ZhaoAAAI 2025 · 35 citations
- The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test AwarenessSahar Abdelnabi, Ahmed SalemNeurIPS 2025 · 28 citations
- Householder Pseudo-Rotation: A Novel Approach to Activation Editing in LLMs with Direction-Magnitude PerspectiveVan-Cuong Pham, Thien NguyenEMNLP 2024
