Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language Models
Chenchen Yuan, Zheyu Zhang, Gjergji Kasneci
摘要
Large language models often display heterogeneous moral preferences across settings. We study inference-time steering toward a desired ethical framework while preserving general competence. We present Convergent-Divergent Routing, which traces and edits minimal branch points inside transformer blocks where ethical-framework-related pathways first converge and then diverge. Gating non-target branches at these loci blocks the downstream propagation while leaving upstream computations intact. We find that this intervention alone increases targeted ethical-framework reasoning. To achieve fine-grained control, we adapt Common Spatial Patterns to the residual stream and extract, for each branch-point layer, a pair of directions that discriminate between utilitarian and deontological frameworks. We then introduce Dual Logit Calibration, a closed-form, minimum--norm update that moves the residual within this two-dimensional subspace so the resulting directional projections align with user-specified preference weights. Experiments on real-life moral dilemmas show that our method reliably achieves preference calibration and largely preserves general capabilities, outperforming recent baselines while providing an interpretable mechanism.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch 等ICLR 2021 · 被引用 878 次
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 被引用 303 次
- Large Language Models are Human-Level Prompt EngineersYongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster 等ICLR 2023 · 被引用 297 次
- ReFT: Representation Finetuning for Language ModelsZhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger 等NeurIPS 2024 · 被引用 233 次
相关 Paper
- PALC: Preference Alignment via Logit CalibrationSanghyun Lee, Hoh InICLR 2026
- Do Morals Guide How LLMs Think? The Role of Ethical Perspectives in General Problem SolvingIseo Kim, Eunjin Hong, Juae KimACL 2026
- Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation FrameworkMohna Chakraborty, Lu Wang, David JurgensEMNLP 2025
- Advancing Automated Ethical Profiling in SE: a Zero-Shot Evaluation of LLM ReasoningPatrizio Migliarini, Mashal Afzal Memon, Marco Autili, Paola InverardiASE 2025
- Moral Alignment for LLM AgentsElizaveta Tennant, Stephen Hailes, Mirco MusolesiICLR 2025
