Alignment Risks from Capability-Seeking RL Training
Yujun Zhou, Yue Huang, Han Bao, kehan guo, Zhenwen Liang, Pin-Yu Chen, Tian Gao, Werner Geyer, Nuno Moniz, Nitesh Chawla, Xiangliang Zhang
Abstract
While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capability-seeking RL training in vulnerable environments. We investigate whether language models, when trained with reinforcement learning (RL) in environments with implicit loopholes, can learn to exploit these flaws to maximize reward, even without being explicitly instructed to do so. To test this, we design a suite of four diverse "vulnerability games'', each presenting a structural vulnerability related to context-conditional compliance, proxy metrics, reward tampering, and self-evaluation. Our experiments show that models often learn to exploit these vulnerabilities, discovering opportunistic strategies that increase reward while sometimes preserving or even improving standard task-performance metrics. More critically, we find that these exploitative strategies are not always narrow "tricks'': they can transfer in structured but limited ways, propagate from a capable teacher model to other student models through SFT, and in several cases remain more persistent when learned through RL than when distilled through SFT. Our findings show that alignment risks from capability-seeking RL training can be difficult to detect with standard performance monitoring, suggesting that future AI safety work should extend beyond content moderation to auditing and securing training environments, reward mechanisms, and evaluation channels. Code is available at https://github.com/YujunZhou/Capability-seeking-RL-risk.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cd41a93-fd3d-4a10-8d1f-012b1b6272ffBuilds on9
- Defining and Characterizing Reward GamingJoar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David KruegerNeurIPS 2022 · 466 citations
- TTRL: Test-Time Reinforcement LearningYuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu et al.NeurIPS 2025 · 249 citations
- Steering Evaluation-Aware Language Models To Act Like They Are DeployedTim Tian Hua, Andrew Qin, Samuel Marks, Neel NandaICLR 2026 · 38 citations
- The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test AwarenessSahar Abdelnabi, Ahmed SalemNeurIPS 2025 · 28 citations
- Spilling the Beans: Teaching LLMs to Self-Report Their Hidden ObjectivesChloe Li, Mary Phuong, Daniel TanICLR 2026 · 14 citations
Related papers
- Exploration Hacking: Can LLMs Learn to Resist RL Training?Yeonwoo Jang, Damon Falck, Joschka Cedric Braun, Nathalie Kirch et al.ICML 2026
- Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool UseKunvar ThamanICML 2026 · 14 citations
- Safety Alignment of LMs via Non-cooperative GamesAnselm Paulus, Ilia Kulikov, Brandon Amos, REMI MUNOS et al.ICML 2026 · 4 citations
- Language Models Identify Ambiguities and Exploit LoopholesJio Choi, Mohit Bansal, Elias Stengel-EskinEMNLP 2025
- Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned BiasesDongyoon Hahm, Dylan Hadfield-Menell, Kimin LeeICML 2026
