Style-Specific Neurons for Steering LLMs in Text Style Transfer
Wen Lai, Viktor Hangya, Alexander Fraser
摘要
Text style transfer (TST) aims to modify the style of a text without altering its original meaning. Large language models (LLMs) demonstrate superior performance across multiple tasks, including TST. However, in zero-shot setups, they tend to directly copy a significant portion of the input text to the output without effectively changing its style. To enhance the stylistic variety and fluency of the text, we present sNeuron-TST, a novel approach for steering LLMs using style-specific neurons in TST. Specifically, we identify neurons associated with the source and target styles and deactivate source-style-only neurons to give target-style words a higher probability, aiming to enhance the stylistic diversity of the generated text. However, we find that this deactivation negatively impacts the fluency of the generated text, which we address by proposing an improved contrastive decoding method that accounts for rapid token probability shifts across layers caused by deactivated source-style neurons. Empirical experiments demonstrate the effectiveness of the proposed method on six benchmarks, encompassing formality, toxicity, politics, politeness, authorship, and sentiment 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Neuron Empirical Gradient: Discovering and Quantifying Neurons' Global Linear ControllabilityXin Zhao, Zehui Jiang, Naoki YoshinagaACL 2025 · 被引用 3 次
- SAEs Are Good for Steering - If You Select the Right FeaturesDana Arad, Aaron Mueller, Yonatan BelinkovEMNLP 2025
- Diff4TST: Masked Diffusion Language Model for Text Style TransferXinchen Ma, Gaole He, Yunshi Lan, Weining QianACL 2026
- Towards a Unified Paradigm of Concept Editing in Large Language ModelsZhuowen Han, Xinwei Wu, Dan Shi, Renren Jin 等EMNLP 2025
它引用的顶会 Paper12
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim 等ICLR 2024 · 被引用 354 次
- ParaDetox: Detoxification with Parallel DataVarvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy 等ACL 2022 · 被引用 96 次
- Contrastive Decoding: Open-ended Text Generation as OptimizationXiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang 等ACL 2023 · 被引用 78 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
相关 Paper
- Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Arbitrary Textual Style Transfer with Small Language ModelsMirac Suzgun, Luke Melas-Kyriazi, Dan JurafskyEMNLP 2022 · 被引用 34 次
- Text Style Transfer Back-TranslationDaimeng Wei, Zhanglin Wu, Hengchao Shang, Zongyao Li 等ACL 2023 · 被引用 10 次
- Masked Based Unsupervised Content TransferRon Mokady, Sagie Benaim, Lior Wolf, Amit BermanoICLR 2020
- SC2: Towards Enhancing Content Preservation and Style Consistency in Long Text Style TransferJie Zhao, Ziyu Guan, Cai Xu, Wei Zhao 等ACL 2024
- Rethinking Style Transformer with Energy-based Interpretation: Adversarial Unsupervised Style Transfer using a Pretrained ModelHojun Cho, Dohee Kim, Seungwoo Ryu, ChaeHun Park 等EMNLP 2022
