Language Models Resist Alignment: Evidence From Data Compression
Jiaming Ji, Kaile Wang, Tianyi Alex Qiu, Boyuan Chen, Jiayi Zhou, Changye Li, Hantao Lou, Josef Dai, Yunhuai Liu, Yaodong Yang
摘要
Large language models (LLMs) may exhibit unintended or undesirable behaviors. Recent works have concentrated on aligning LLMs to mitigate harmful outputs. Despite these efforts, some anomalies indicate that even a wellconducted alignment process can be easily circumvented, whether intentionally or accidentally. Does alignment fine-tuning yield have robust effects on models, or are its impacts merely superficial? In this work, we make the first exploration of this phenomenon from both theoretical and empirical perspectives. Empirically, we demonstrate the elasticity of postalignment models, i.e., the tendency to revert to the behavior distribution formed during the pretraining phase upon further fine-tuning. Leveraging compression theory, we formally deduce that fine-tuning disproportionately undermines alignment relative to pre-training, potentially by orders of magnitude. We validate the presence of elasticity through experiments on models of varying types and scales. Specifically, we find that model performance declines rapidly before reverting to the pre-training distribution, after which the rate of decline drops significantly. Furthermore, we further reveal that elasticity positively correlates with the increased model size and the expansion of pre-training data. Our findings underscore the need to address the inherent elasticity of LLMs to mitigate their resistance to alignment. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained LearningBorong Zhang, Yuhao Zhang, Jiaming Ji, Yingshan Lei 等NeurIPS 2025 · 被引用 84 次
- SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMsJiacheng Lin, Zhongruo Wang, Kun Qian, Tian Wang 等ICLR 2026 · 被引用 25 次
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentCameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim 等ICML 2026 · 被引用 22 次
- Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning DatasetsLei Hsiung, Tianyu Pang, Yung-Chen Tang, Linyue Song 等ACL 2026 · 被引用 22 次
- FEEDBACK FRICTION: LLMs Struggle to Fully Incorporate External FeedbackDongwei Jiang, Bowei Zhang, Andrew Wang, Nicholas Andrews 等NeurIPS 2025 · 被引用 11 次
它引用的顶会 Paper14
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji 等ICLR 2024 · 被引用 656 次
相关 Paper
- Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update AlignmentYavuz Faruk Bakman, Duygu Nur Yaldiz, Eleni Triantafillou, Peter Kairouz 等ICML 2026 · 被引用 2 次
- Alleviating the Fear of Losing Alignment in LLM Fine-tuningKang Yang, Guanhong Tao, Xun Chen, Jun XuS&P 2025
- Understanding the Learning Dynamics of Alignment with Human FeedbackShawn Im, Yixuan LiICML 2024 · 被引用 18 次
- Safety Alignment Should be Made More Than Just a Few Tokens DeepXiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma 等ICLR 2025
- Emergent Misalignment is Easy, Narrow Misalignment is HardAnna Soligo, Edward Turner, Senthooran Rajamanoharan, Neel NandaICLR 2026 · 被引用 25 次
