Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization
Weixu Zhang, Ye Yuan, Changjiang Han, Yuxing Tian, Zipeng Sun, Linfeng Du, Jikun Kang, Hong Kang, Xue Liu, Haolun Wu
Abstract
Large Language Models (LLMs) exhibit strong implicit personalization ability, yet most existing approaches treat this behavior as a black box, relying on prompt engineering or fine tuning on user data. In this work, we adopt a mechanistic interpretability perspective and hypothesize the existence of a sparse set of Preference Heads, attention heads that encode user specific stylistic and topical preferences and exert a causal influence on generation. We introduce Differential Preference Steering (DPS), a training free framework that (1) identifies Preference Heads through causal masking analysis and (2) leverages them for controllable and interpretable personalization at inference time. DPS computes a Preference Contribution Score (PCS) for each attention head, directly measuring its causal impact on user aligned outputs. During decoding, we contrast model predictions with and without Preference Heads, amplifying the difference between personalized and generic logits to selectively strengthen preference aligned continuations. Experiments on widely used personalization benchmarks across multiple LLMs demonstrate consistent gains in personalization fidelity while preserving content coherence and low computational overhead. Beyond empirical improvements, DPS provides a mechanistic explanation of where and how personalization emerges within transformer architectures. Our implementation is publicly available. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 717edab2-e06c-41db-bff6-0265126b2ad5Cited by top-tier papers1
Ask how each one uses itBuilds on10
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim et al.ICLR 2024 · 354 citations
- Iterative Reasoning Preference OptimizationRichard Yuanzhe Pang, Weizhe Yuan, He He, Kyunghyun Cho et al.NeurIPS 2024 · 287 citations
- Aligning LLM Agents by Learning Latent Preference from User EditsGe Gao, Alexey Taymanov, Eduardo Salinas, Paul Mineiro et al.NeurIPS 2024 · 102 citations
- Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target AtomsMengru Wang, Ziwen Xu, Shengyu Mao, Shumin Deng et al.ACL 2025 · 19 citations
- OpenDecoder: Open Large Language Model Decoding to Incorporate Document Quality in RAGFengran Mo, Zhan Su, Yuchen Hui, Jinghan Zhang et al.WWW 2026 · 8 citations
Related papers
- NextQuill: Causal Preference Modeling for Enhancing LLM PersonalizationXiaoyan Zhao, Juntao You, Yang Zhang, Wenjie Wang et al.ICLR 2026 · 38 citations
- Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference OptimizationYuanpu Cao, Tianrong Zhang, Bochuan Cao, Ziyi Yin et al.NeurIPS 2024 · 135 citations
- Learning Personalized Alignment for Evaluating Open-ended Text GenerationDanqing Wang, Kevin Yang, Hanlin Zhu, Xiaomeng Yang et al.EMNLP 2024 · 2 citations
- Causally Motivated Sycophancy Mitigation for Large Language ModelsHaoxi Li, Xueyang Tang, Jie Zhang, Song Guo et al.ICLR 2025
- Steer Like the LLM: Activation Steering that Mimics PromptingGeert Heyman, Frederik VandeputteICML 2026
