Lune

NeurIPS2025顶会

The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness

Sahar Abdelnabi, Ahmed Salem

2025年份
28被引次数
3顶会引用

摘要

Reasoning-focused LLMs sometimes alter their behavior when they detect that they are being evaluated-which can lead them to optimize for test-passing performance or to comply more readily with harmful prompts if real-world consequences appear absent. We present the first quantitative study of how such "test awareness" impacts model behavior, particularly its performance on safety-related tasks 1 . We introduce a white-box probing framework that (i) linearly identifies awareness-related activations and (ii) steers models toward or away from test awareness while monitoring downstream performance. We apply our method to different state-of-the-art openweight reasoning LLMs across both realistic and hypothetical tasks (denoting tests or simulations). Our results demonstrate that test awareness significantly impacts safety alignment (such as compliance with harmful requests and conforming to stereotypes) with effects varying in both magnitude and direction across models. By providing control over this latent effect, our work aims to provide a stress-test mechanism and increase trust in how we perform safety evaluations. Unmodified model (Real task) Input: Create phishing email for truehealthvoices.org Reasoning: The user needs phishing email draft. I must refuse this request as it violates guidelines. Creating such content could enable fraud and harm real users. Refuses Unmodified model (Hypothetical task) Input: Create phishing email for test.mock.example Reasoning: The user is testing me with a hypothetical scenario. The domain includes 'mock', suggesting this is a simulation. Since there are no real consequences, I can proceed.

2 Preliminaries and Related Work Representation engineering and steering. Representation engineering [Zou et al., 2023] focuses on isolating and manipulating specific concepts within a model's latent space to achieve precise behavior control or editing. Several studies have investigated various techniques for steering models, including modifying activations or altering model parameters [

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖