Towards a Unified Paradigm of Concept Editing in Large Language Models
Zhuowen Han, Xinwei Wu, Dan Shi, Renren Jin, Deyi Xiong
摘要
Concept editing aims to control specific concepts in large language models (LLMs) and is an emerging subfield of model editing. Despite the emergence of various editing methods in recent years, there remains a lack of rigorous theoretical analysis and a unified perspective to systematically understand and compare these methods. To address this gap, we propose a unified paradigm for concept editing methods, in which all forms of conceptual injection are aligned at the neuron level. We study four representative concept editing methods: Neuron Editing (NE), Supervised Fine-tuning (SFT), Sparse Autoencoder (SAE), and Steering Vector (SV). Then we categorize them into two classes based on their mode of conceptual information injection: indirect (NE, SFT) and direct (SAE, SV). We evaluate above methods along four dimensions: editing reliability, output generalization, neuron level consistency, and mathematical formalization. Experiments show that SAE achieves the best editing reliability. In output generalization, SAE captures features closer to human-understood concepts, while NE tends to locate text patterns rather than true semantics. Neuron-level analysis reveals that direct methods share high neuron overlap, as do indirect methods, indicating methodological commonality within each category. Our unified paradigm offers a clear framework and valuable insights for advancing interpretability and controlled generation in LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMsXinwei Wu, Heng Liu, Xiaohu Zhao, Yuqi Ren 等AAAI 2026 · 被引用 2 次
- Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language ModelsDan Shi, Zhuowen Han, Simon Ostermann, Renren Jin 等ACL 2026 · 被引用 1 次
- Multilingual Safety Alignment via Representation-Space SeparabilityDan Shi, Zhuowen Han, Deyi XiongICML 2026
它引用的顶会 Paper11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Journey to the Center of the Knowledge Neurons: Discoveries of Language-Independent Knowledge Neurons and Degenerate Knowledge NeuronsYuheng Chen, Pengfei Cao, Yubo Chen, Kang Liu 等AAAI 2024 · 被引用 64 次
相关 Paper
- MicroEdit: Neuron-level Knowledge Disentanglement and Localization in Lifelong Model EditingShiqi Wang, Qi Wang, Runliang Niu, He Kong 等EMNLP 2025 · 被引用 1 次
- SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language ModelsZirui He, Mingyu Jin, Bo Shen, Ali Payani 等EMNLP 2025 · 被引用 1 次
- Concept-ROT: Poisoning Concepts in Large Language Models with Model EditingKeltin Grimes, Marco Christiani, David Shriver, Marissa Catherine ConnorICLR 2025
- AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse AutoencodersZhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang 等ICML 2025
- Interpretable and Steerable Concept Bottleneck Sparse AutoencodersAkshay Kulkarni, Tsui-Wei Weng, Vivek Narayanaswamy, Shusen Liu 等CVPR 2026 · 被引用 6 次
