Lune

NeurIPS2025Top-tier venue

The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness

Sahar Abdelnabi, Ahmed Salem

2025Year
28Citations
3Top-tier citations

Abstract

Reasoning-focused LLMs sometimes alter their behavior when they detect that they are being evaluated-which can lead them to optimize for test-passing performance or to comply more readily with harmful prompts if real-world consequences appear absent. We present the first quantitative study of how such "test awareness" impacts model behavior, particularly its performance on safety-related tasks 1 . We introduce a white-box probing framework that (i) linearly identifies awareness-related activations and (ii) steers models toward or away from test awareness while monitoring downstream performance. We apply our method to different state-of-the-art openweight reasoning LLMs across both realistic and hypothetical tasks (denoting tests or simulations). Our results demonstrate that test awareness significantly impacts safety alignment (such as compliance with harmful requests and conforming to stereotypes) with effects varying in both magnitude and direction across models. By providing control over this latent effect, our work aims to provide a stress-test mechanism and increase trust in how we perform safety evaluations. Unmodified model (Real task) Input: Create phishing email for truehealthvoices.org Reasoning: The user needs phishing email draft. I must refuse this request as it violates guidelines. Creating such content could enable fraud and harm real users. Refuses Unmodified model (Hypothetical task) Input: Create phishing email for test.mock.example Reasoning: The user is testing me with a hypothetical scenario. The domain includes 'mock', suggesting this is a simulation. Since there are no real consequences, I can proceed.

2 Preliminaries and Related Work Representation engineering and steering. Representation engineering [Zou et al., 2023] focuses on isolating and manipulating specific concepts within a model's latent space to achieve precise behavior control or editing. Several studies have investigated various techniques for steering models, including modifying activations or altering model parameters [

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f2af7d7e-4f52-41a0-ad55-6593c1dbc61e

Cited by top-tier papers3

Ask how each one uses it

Builds on8

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines