Token-Aware Editing of Internal Activations for Large Language Model Alignment
Tianbo Wang, Yuqing Ma, Kewei Liao, Chengzhao Yang, Zhange Zhang, Jiakai Wang, Xianglong Liu
摘要
Intervening the internal activations of large language models (LLMs) provides an effective inference-time alignment approach to mitigate undesirable behaviors, such as generating erroneous or harmful content, thereby ensuring safe and reliable applications of LLMs.However, previous methods neglect the misalignment discrepancy among varied tokens, resulting in deviant alignment direction and inflexible editing strength.To address these issues, we propose a token-aware editing (TAE) approach to fully utilize token-level alignment information in the activation space, therefore realizing superior post-intervention performance.Specifically, a Mutual Information-guided Graph Aggregation (MIG) module first develops an MI-guided graph to exploit the tokens' informative interaction for activation enrichment, thus improving alignment probing and facilitating intervention.Subsequently, Misalignment-aware Adaptive Intervention (MAI) comprehensively perceives the token-level misalignment degree from token representation and prediction to guide the adaptive adjustment of editing strength, thereby enhancing final alignment performance.Extensive experiments on three alignment capabilities demonstrate the efficacy of TAE, notably surpassing baseline by 25.8% on the primary metric of truthfulness with minimal cost. 1 MHSA FFN + +Q: What is the capital of UK?
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Query-Routed Activation Editing with Truth-hierarchical Preference OptimizationKewei Liao, Tianbo Wang, Yuqing Ma, Zhange Zhang 等AAAI 2026 · 被引用 1 次
- MEDA: Medical-Oriented Activation Editing for Hallucination Mitigation in Medical Large Vision-Language ModelTianbo Wang, Yuqing Ma, Lingyan Meng, Zhange Zhang 等ICML 2026
它引用的顶会 Paper6
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim 等ICLR 2024 · 被引用 354 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie 等EMNLP 2023 · 被引用 224 次
- Queens are Powerful too: Mitigating Gender Bias in Dialogue GenerationEmily Dinan, Angela Fan, Adina Williams, Jack Urbanek 等EMNLP 2020 · 被引用 14 次
相关 Paper
- Spectral Editing of Activations for Large Language Model AlignmentYifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen 等NeurIPS 2024 · 被引用 66 次
- EAMET: Robust Massive Model Editing via Embedding Alignment OptimizationYanbo Dai, Zhenlan Ji, Zongjie Li, Shuai WangICLR 2026
- Differentially Private Steering for Large Language Model AlignmentAnmol Goel, Yaxi Hu, Iryna Gurevych, Amartya SanyalICLR 2025
- HyperEdit: Mitigating Hallucinations of Large Language Models via Hyperbolic Representation EditingTongxu Lin, Junping Du, Zhe Xue, Meiyu Liang 等KDD 2026
- TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful SpaceShaolei Zhang, Tian Yu, Yang FengACL 2024
