Multi-Adapter Representation Interventions via Energy Calibration
Manjiang Yu, Hongji Li, Junwei Chen, Xue Li, Priyanka Singh, YANG CAO, Lijie Hu
摘要
Representation intervention has emerged as a promising paradigm for aligning large language models toward desired behaviors without modifying model weights. Existing methods typically apply a fixed intervention uniformly across all inputs. However, we find that the appropriate intervention direction and strength vary substantially across samples, and such indiscriminate intervention leads to degradation of general capabilities on benign inputs. To address these challenges, we propose Multi-Adapter Representation Interventions via Energy Calibration (MARI). Specifically, we introduce a competitive multi-adapter mechanism in which specialized experts capture non-linear correction patterns and adaptively determine the appropriate intervention direction and strength for different samples. Furthermore, we design an energy-based gating module that leverages internal propagation dynamics to distinguish inputs that are applicable for intervention. Extensive experiments across diverse model families and parameter scales demonstrate that MARI achieves state-of-the-art alignment performance. Our method significantly improves performance on TruthfulQA, BBQ, and safety benchmarks, while maintaining and even improving general capabilities on tasks such as MMLU and ARC. Our code is available at https://github.com/V1centNevwake/MARI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
相关 Paper
- Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language ModelsAnirudh Sundar, Sinead Williamson, Katherine Metcalf, Barry-John Theobald 等ACL 2025 · 被引用 8 次
- Concept Concentration for Faithful Representation InterventionHongzheng Yang, Yongqiang Chen, Zeyu Qin, Tongliang Liu 等ICML 2026
- Steering When Necessary: Flexible Steering Large Language Models with BacktrackingZifeng Cheng, Jinwei Gan, Zhiwei Jiang, Cong Wang 等NeurIPS 2025 · 被引用 9 次
- Multi-Attribute Steering of Language Models via Targeted InterventionDuy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit BansalACL 2025 · 被引用 30 次
- Learning "Partner-Aware" Collaborators in Multi-Party CollaborationAbhijnan Nath, Nikhil KrishnaswamyNeurIPS 2025 · 被引用 2 次
