Consistency Training Can Entrench Misalignment
David Africa, Arathi Mani
摘要
Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures. Such methods are simple, scalable, and largely label-free, but their effects on model alignment remain poorly understood. Could the self-bootstrapping nature of these methods amplify undesired behavior in models? We test seven consistency training methods on 108 "model organisms 1 ": open-source models (7B-70B) fine-tuned to exhibit various forms of controlled misaligned behavior. We find that outcomes vary significantly: consistency training generally suppresses reward hacking and emergent misalignment but amplifies sycophancy. We present evidence that distribution shifts induced by the consistency labeling process, rather than variation in the selection operators, may be the primary driver of systematic alignment effects. Finally, we present a unifying theoretical framework to derive conditions under which consistency training will amplify or suppress misalignment. In total, our study establishes that consistency training is not alignment-neutral, and that its use in critical systems should be carefully audited.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
相关 Paper
- Self-Consistency Preference OptimizationArchiki Prasad, Weizhe Yuan, Richard Yuanzhe Pang, Jing Xu 等ICML 2025
- ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-TrainingYu Liang, Liangxin Liu, Longzheng Wang, Yan Wang 等ACL 2026
- When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming LoopYang Zhang, Xiukun Wei, Xueru ZhangICML 2026
- Removing Sandbagging in LLMs by Training with Weak SupervisionEmil Ryd, Henning Bartsch, Julian Stastny, Joe Benton 等ICML 2026 · 被引用 1 次
- One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward ModelsDaniel Fein, Max Lamparth, Violet Xiang, Mykel Kochenderfer 等ICML 2026
