The Realignment Problem: When Right becomes Wrong in LLMs
Aakash Sen Sharma, Debdeep Sanyal, Manodeep Ray, Vivek Srivastava, Shirish Karande, Murari Mandal
摘要
Post-training alignment of large language models (LLMs) relies on large-scale human annotations guided by policy specifications that change over time. Cultural shifts, value reinterpretations, and regulatory or industrial updates make static alignment increasingly brittle. As policies evolve, deployed models can diverge from current alignment objectives, creating an Alignment–Reality Gap that is difficult to audit or correct. Existing remediation typically requires re-annotation under revised guidelines, which introduces systematic challenges, including guideline ambiguity, annotator interpretation drift, and reduced consistency at scale. We introduce TRACE (Triage and Re-align by Alignment Conflict Evaluation), a framework that transforms re-alignment into a structured optimization problem over existing data without requiring fresh human annotation. Leveraging a stronger model as a proxy judge, TRACE operates via a three-stage pipeline: (1) triaging preference pairs into inversion, suppression, or retention categories based on alignment conflicts; (2) computing an alignment impact score via bi-level optimization to prioritize high-leverage samples; and (3) executing updates using a hybrid objective that combines relational losses (e.g., IPO) for preference inversion and punitive losses (e.g., NPO) for response suppression. Experiments on Qwen2.5-7B, Gemma-2-9B, and Llama-3.1-8B demonstrate robust re-alignment on synthetic benchmarks and the PKU-SafeRLHF dataset without degrading general utility. This work provides a scalable approach for LLM realignment under evolving data annotation policies and alignment guidelines. We release our code here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 被引用 365 次
- Provably Robust DPO: Aligning Language Models with Noisy FeedbackSayak Ray Chowdhury, Anush Kini, Nagarajan NatarajanICML 2024 · 被引用 118 次
- Influence Functions in Deep Learning Are FragileSamyadeep Basu, Phillip Pope, Soheil FeiziICLR 2021 · 被引用 15 次
相关 Paper
- Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual FeedbackYafu Li, Xuyang Hu, Xiaoye Qu, Linjie Li 等ICML 2025
- Keep the Best, Forget the Rest: Reliable Alignment with Order-Aware Preference OptimizationJiahui Zhu, Yuanjie Shi, Xiyue Peng, Xin Liu 等ICLR 2026
- Conflict-Aware Adaptive Alignment for LLM Hallucination MitigationRuohan Zong, Yang Zhang, WangICML 2026
- MetaAligner: Towards Generalizable Multi-Objective Alignment of Language ModelsKailai Yang, Zhiwei Liu, Qianqian Xie, Jimin Huang 等NeurIPS 2024 · 被引用 49 次
- Governance in Motion: Co-evolution of Constitutions and AI models for Scalable SafetyChenhao Huang, Ziyu Shen, Yicong Ren, Huiyuan Zheng 等EMNLP 2025
