Householder Pseudo-Rotation: A Novel Approach to Activation Editing in LLMs with Direction-Magnitude Perspective
Van-Cuong Pham, Thien Nguyen
摘要
Activation Editing, which involves directly editing the internal representations of large language models (LLMs) to alter their behaviors and achieve desired properties, has emerged as a promising area of research. Existing works primarily treat LLMs' activations as points in space and modify them by adding steering vectors. However, this approach is limited in its ability to achieve greater performance improvement while maintaining the necessary consistency of activation magnitudes. To overcome these issues, we propose a novel editing method that views activations in terms of their directions and magnitudes. Our method, named Householder Pseudo-Rotation (HPR), mimics the rotation transformation, thus preserving activation norms and resulting in an improved performance on various safety benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Angular Steering: Behavior Control via Rotation in Activation SpaceMinh Hieu Vu, Tan M. NguyenNeurIPS 2025 · 被引用 53 次
- Multi-Attribute Steering of Language Models via Targeted InterventionDuy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit BansalACL 2025 · 被引用 30 次
- ODESteer: A Unified ODE-Based Steering Framework for LLM AlignmentHongjue Zhao, Haosen Sun, Jiangtao Kong, Xiaochang Li 等ICLR 2026 · 被引用 16 次
- GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMsDuy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit BansalACL 2026 · 被引用 4 次
- Query-Routed Activation Editing with Truth-hierarchical Preference OptimizationKewei Liao, Tianbo Wang, Yuqing Ma, Zhange Zhang 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
相关 Paper
- SAME: Safety-Aware Model Editing Guided by Safety TransformationJiayi Wang, Shipeng Wang, Ji Wu, Jian SunACL 2026
- Activation Steering with a Feedback ControllerDung Viet Nguyen, Yen Nhi Pham, Hieu M. Vu, Lei Zhang 等ICLR 2026 · 被引用 13 次
- Differentially Private Steering for Large Language Model AlignmentAnmol Goel, Yaxi Hu, Iryna Gurevych, Amartya SanyalICLR 2025
- Spectral Editing of Activations for Large Language Model AlignmentYifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen 等NeurIPS 2024 · 被引用 66 次
- Steering Evaluation-Aware Language Models To Act Like They Are DeployedTim Tian Hua, Andrew Qin, Samuel Marks, Neel NandaICLR 2026 · 被引用 38 次
