Introspection Adapters: Training LLMs to Report Their Learned Behaviors
Keshav Shenoy, Li Yang, Abhay Sheshadri, Jack Lindsey, Samuel Marks, Rowan Wang
摘要
Can we train LLMs to introspect, i.e. to faithfully describe their own behaviors in natural language? Prior work has shown some, limited, success. However, it is difficult to scale introspection training due to a lack of ground-truth labels. In this work, we study an approach to introspection training which side-steps this data bottleneck. Given a target model , our method works by fine-tuning models from with implanted behaviors (such as downplaying medical problems); the pairs serve as labeled introspection training data. We then train an introspection adapter (IA): a LoRA adapter jointly optimized across the fine-tunes which causes them to verbalize their implanted behaviors. This IA induces faithful introspection in fine-tunes of that were trained in very different ways from the , as well as in itself. This is surprising because the IA was never trained on . To demonstrate the utility of IAs, we use them to successfully audit misaligned models introduced in prior work. IAs can also be used to detect fine-tuning API attacks which train models to comply with encrypted harmful requests. Notably, IAs are more effective when applied to larger models. Overall, our results suggest that IAs are a scalable, effective, and practically useful approach to LLM introspection training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
相关 Paper
- Split Personality Training: Revealing Latent Knowledge Through Alternate PersonalitiesFlorian Dietz, William Wale, Oscar Gilg, Robert McCarthy 等ICML 2026 · 被引用 2 次
- Causal-Guided Detoxify Backdoor Attack of Open-Weight LoRA ModelsLinzhi Chen, Yang Sun, Hongru Wei, Yuqi ChenNDSS 2026 · 被引用 4 次
- In-Training Defenses Against Emergent Misalignment in Language ModelsDavid Kaczér, Magnus Jørgenvåg, Clemens Vetter, Esha Afzal 等ICML 2026 · 被引用 13 次
- Looking Inward: Language Models Can Learn About Themselves by IntrospectionFelix Jedidja Binder, James Chua, Tomek Korbak, Henry Sleight 等ICLR 2025 · 被引用 5 次
- Finding and Reactivating Post-Trained LLMs' Hidden Safety MechanismsMingjie Li, Wai Man Si, Michael Backes, Yang Zhang 等NeurIPS 2025 · 被引用 4 次
