Stress-Testing Capability Elicitation With Password-Locked Models
Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, David Krueger
Abstract
To determine the safety of large language models (LLMs), AI developers must be able to assess their dangerous capabilities. But simple prompting strategies often fail to elicit an LLM's full capabilities. One way to elicit capabilities more robustly is to fine-tune the LLM to complete the task. In this paper, we investigate the conditions under which fine-tuning-based elicitation suffices to elicit capabilities. To do this, we introduce password-locked models, LLMs fine-tuned such that some of their capabilities are deliberately hidden. Specifically, these LLMs are trained to exhibit these capabilities only when a password is present in the prompt, and to imitate a much weaker LLM otherwise. Password-locked models enable a novel method of evaluating capabilities elicitation methods, by testing whether these password-locked capabilities can be elicited without using the password. We find that a few high-quality demonstrations are often sufficient to fully elicit password-locked capabilities. More surprisingly, fine-tuning can elicit other capabilities that have been locked using the same password, or even different passwords. Furthermore, when only evaluations, and not demonstrations, are available, approaches like reinforcement learning are still often able to elicit capabilities. Overall, our findings suggest that fine-tuning is an effective method of eliciting hidden capabilities of current models, but may be unreliable when high-quality demonstrations are not available, e.g. as may be the case when models' (hidden) capabilities exceed those of human demonstrators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f656a89-c48f-46ca-9e1a-ff89fae28992Cited by top-tier papers12
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMsKyle O'Brien, Stephen Casper, Quentin Anthony, Tomek Korbak et al.ICLR 2026 · 59 citations
- Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept SpaceCore Francisco Park, Maya Okawa, Andrew Lee, Ekdeep Singh Lubana et al.NeurIPS 2024 · 39 citations
- Noise Injection Reveals Hidden Capabilities of Sandbagging Language ModelsCameron Tice, Philipp Alexander Kreer, Nathan Helm-Burger, Prithviraj Singh Shahani et al.NeurIPS 2025 · 24 citations
- CTRL-ALT-DECEIT Sabotage Evaluations for Automated AI R&DFrancis Ward, Teun van der Weij, Hanna Gábor, Sam Martin et al.NeurIPS 2025 · 11 citations
- The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent HomologyAideen Fay, Inés García-Redondo, Qiquan Wang, Haim Dubossarsky et al.ICLR 2026 · 8 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
Related papers
- The Elicitation Game: Evaluating Capability Elicitation TechniquesFelix Hofstätter, Teun van der Weij, Jayden Teoh, Rada Djoneva et al.ICML 2025
- Quantifying Elicitation of Latent Capabilities in Language ModelsElizabeth Donoway, Hailey Joren, Arushi Somani, Henry Sleight et al.NeurIPS 2025 · 4 citations
- AI Sandbagging: Language Models can Strategically Underperform on EvaluationsTeun van der Weij, Felix Hofstätter, Oliver Jaffe, Samuel F. Brown et al.ICLR 2025
- Exploration Hacking: Can LLMs Learn to Resist RL Training?Yeonwoo Jang, Damon Falck, Joschka Cedric Braun, Nathalie Kirch et al.ICML 2026
- Can Foundation LLMs Accurately Estimate Password Strength and Provide Appropriate Password Feedback?Madison Pickering, Garrison Hinson-Hasty, Luca Dovichi, Helena Williams et al.S&P 2026
