Lune

ICML2026Top-tier venue

Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models

Jongwook Han, Jongwon Lim, Injin Kong, Yohan Jo

2026Year

Abstract

Large language models can express values in two main ways: (1) intrinsic\textit{intrinsic} expression, reflecting the model's inherent values learned during training, and (2) prompted\textit{prompted} expression, elicited by explicit prompts. Given their widespread use in value alignment, it is paramount to clearly understand their underlying mechanisms, particularly whether they mostly overlap (as one might expect) or rely on distinct mechanisms. We analyze this largely understudied problem at the mechanistic level using two approaches: (1) value vectors\textit{value vectors}, feature directions representing value mechanisms extracted from the residual stream, and (2) value neurons\textit{value neurons}, MLP neurons that contribute to value vectors. We demonstrate that intrinsic and prompted value mechanisms partly share common components crucial for inducing value expression, generalizing across languages and reconstructing theoretical inter-value correlations in the model's internal representations. Yet, each mechanism also possesses unique components that fulfill distinct roles. In particular, the intrinsic mechanism activates in more diverse value-related scenarios and promotes response diversity, whereas the prompted mechanism strengthens instruction compliance, taking effect even in distant tasks like jailbreaking.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b6268031-af79-41dc-87a6-877c07dfe61f

Builds on29

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines