StyliTruth : Unlocking Stylized yet Truthful LLM Generation via Disentangled Steering
Chenglei Shen, Zhongxiang Sun, Teng Shi, Xiao Zhang, Jun Xu
摘要
Generating stylized large language model (LLM) responses via representation editing is a promising way for fine-grained output control. However, there exists an inherent trade-off: imposing a distinctive style often degrades truthfulness. Existing representation editing methods, by naively injecting style signals, overlook this collateral impact and frequently contaminate the model’s core truthfulness representations, resulting in reduced answer correctness. We term this phenomenon stylization-induced truthfulness collapse. We attribute this issue to latent coupling between style and truth directions in certain key attention heads, and propose StyliTruth, a mechanism that preserves stylization while keeping truthfulness intact. StyliTruth separates the style-relevant and truth-relevant subspaces in the model’s representation space via an orthogonal deflation process. This decomposition enables independent control of style and truth in their own subspaces, minimizing interference. By designing adaptive, token-level steering vectors within each subspace, we dynamically and precisely control the generation process to maintain both stylistic fidelity and truthfulness. We validate our method on multiple styles and languages. Extensive experiments and analyses show that StyliTruth significantly reduces stylization-induced truthfulness collapse and outperforms existing inference-time intervention methods in balancing style adherence with truthfulness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Model Tells You What to Discard: Adaptive KV Cache Compression for LLMsSuyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang 等ICLR 2024 · 被引用 432 次
- Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsGuillermo Ortiz-Jiménez, Alessandro Favero, Pascal FrossardNeurIPS 2023 · 被引用 272 次
- Controlled Decoding from Language ModelsSidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li 等ICML 2024 · 被引用 130 次
- Aligning Large Language Models with Representation Editing: A Control PerspectiveLingkai Kong, Haorui Wang, Wenhao Mu, Yuanqi Du 等NeurIPS 2024 · 被引用 80 次
相关 Paper
- DRESSing Up LLM: Efficient Stylized Question-Answering via Style Subspace EditingXinyu Ma, Yifeng Xu, Yang Lin, Tianlong Wang 等ICLR 2025
- TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful SpaceShaolei Zhang, Tian Yu, Yang FengACL 2024
- HyperEdit: Mitigating Hallucinations of Large Language Models via Hyperbolic Representation EditingTongxu Lin, Junping Du, Zhe Xue, Meiyu Liang 等KDD 2026
- Neural Stylistic Response Generation with Disentangled Latent VariablesQingfu Zhu, Wei-Nan Zhang, Ting Liu, William Yang WangACL 2021
- FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language ModelsZixuan Weng, Jinghuai Zhang, Kunlin Cai, Ying Li 等ACL 2026
