Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models
Jongwook Han, Jongwon Lim, Injin Kong, Yohan Jo
Abstract
Large language models can express values in two main ways: (1) expression, reflecting the model's inherent values learned during training, and (2) expression, elicited by explicit prompts. Given their widespread use in value alignment, it is paramount to clearly understand their underlying mechanisms, particularly whether they mostly overlap (as one might expect) or rely on distinct mechanisms. We analyze this largely understudied problem at the mechanistic level using two approaches: (1) , feature directions representing value mechanisms extracted from the residual stream, and (2) , MLP neurons that contribute to value vectors. We demonstrate that intrinsic and prompted value mechanisms partly share common components crucial for inducing value expression, generalizing across languages and reconstructing theoretical inter-value correlations in the model's internal representations. Yet, each mechanism also possesses unique components that fulfill distinct roles. In particular, the intrinsic mechanism activates in more diverse value-related scenarios and promotes response diversity, whereas the prompted mechanism strengthens instruction compliance, taking effect even in distant tasks like jailbreaking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6268031-af79-41dc-87a6-877c07dfe61fBuilds on29
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
Related papers
- Do LLMs have Consistent Values?Naama Rozen, Liat Bezalel, Gal Elidan, Amir Globerson et al.ICLR 2025
- Inertia in Moral and Value Judgments of Large Language ModelsBruce W. Lee, Yeongheon Lee, Hyunsoo ChoACL 2026 · 5 citations
- SOM Directions Are Better than One: Multi-Directional Refusal Suppression in Language ModelsGiorgio Piras, Raffaele Mura, Fabio Brau, Luca Oneto et al.AAAI 2026 · 4 citations
- Internal Value Alignment in Large Language Models through Controlled Value Vector ActivationHaoran Jin, Meng Li, Xiting Wang, Zhihao Xu et al.ACL 2025
- Reward Models Inherit Value Biases from PretrainingBrian R. Christian, Jessica A. F. Thompson, Elle Michelle Yang, Vincent Adam et al.ICLR 2026 · 4 citations
