Steering Evaluation-Aware Language Models To Act Like They Are Deployed
Tim Tian Hua, Andrew Qin, Samuel Marks, Neel Nanda
摘要
Large language models (LLMs) can sometimes detect when they are being evaluated and adjust their behavior to appear more aligned, compromising the reliability of safety evaluations. In this paper, we show that adding a steering vector to an LLM's activations can suppress evaluation-awareness and make the model act like it is deployed during evaluation. To study our steering technique, we train an LLM to exhibit evaluation-aware behavior using a two-step training process designed to mimic how this behavior could emerge naturally. First, we perform continued pretraining on two sets of documents describing its behavior. The first says that our model uses Python type hints during evaluation but not during deployment. The second says that our model can recognize that the presence of a certain evaluation cue always means that it is being tested. Then, we train the model with expert iteration to use Python type hints in evaluation settings. The resulting model is evaluation-aware: it writes type hints in evaluation contexts more than deployment contexts. We find that activation steering can suppress evaluation awareness and make the model behave during evaluation as it would during deployment. Importantly, we constructed our steering vector using the original model before our additional training. Our results suggest that AI evaluators could improve the reliability of safety evaluations by steering models to act like they are deployed. * Equal contribution. † Joint supervision, order randomized.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- How do LLMs Compute Verbal Confidence?Dharshan Kumaran, Arthur Conmy, Federico Barbero, Simon Osindero 等ICML 2026 · 被引用 16 次
- Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood ConsistencyHaoming Xu, Ningyuan Zhao, Yunzhi Yao, Weihong Xu 等ACL 2026 · 被引用 2 次
- Alignment Risks from Capability-Seeking RL TrainingYujun Zhou, Yue Huang, Han Bao, kehan guo 等ICML 2026
- Sycophancy Towards Researchers Drives Performative MisalignmentDavid Baek, Xinnuo Li, Anay Gupta, Taslim Mahbub 等ICML 2026
它引用的顶会 Paper8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 被引用 296 次
相关 Paper
- Activation Steering with a Feedback ControllerDung Viet Nguyen, Yen Nhi Pham, Hieu M. Vu, Lei Zhang 等ICLR 2026 · 被引用 13 次
- Endogenous Resistance to Activation Steering in Language ModelsAlex McKenzie, Keenan Pepper, Stijn Servaes, Martin Leitgab 等ICML 2026 · 被引用 3 次
- SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation SteeringZouying Cao, Yifei Yang, Hai ZhaoAAAI 2025 · 被引用 35 次
- The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test AwarenessSahar Abdelnabi, Ahmed SalemNeurIPS 2025 · 被引用 28 次
- Householder Pseudo-Rotation: A Novel Approach to Activation Editing in LLMs with Direction-Magnitude PerspectiveVan-Cuong Pham, Thien NguyenEMNLP 2024
