Style-Specific Neurons for Steering LLMs in Text Style Transfer
Wen Lai, Viktor Hangya, Alexander Fraser
Abstract
Text style transfer (TST) aims to modify the style of a text without altering its original meaning. Large language models (LLMs) demonstrate superior performance across multiple tasks, including TST. However, in zero-shot setups, they tend to directly copy a significant portion of the input text to the output without effectively changing its style. To enhance the stylistic variety and fluency of the text, we present sNeuron-TST, a novel approach for steering LLMs using style-specific neurons in TST. Specifically, we identify neurons associated with the source and target styles and deactivate source-style-only neurons to give target-style words a higher probability, aiming to enhance the stylistic diversity of the generated text. However, we find that this deactivation negatively impacts the fluency of the generated text, which we address by proposing an improved contrastive decoding method that accounts for rapid token probability shifts across layers caused by deactivated source-style neurons. Empirical experiments demonstrate the effectiveness of the proposed method on six benchmarks, encompassing formality, toxicity, politics, politeness, authorship, and sentiment 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c245ef73-7f89-4f9d-bd98-e5c94dbb3314Cited by top-tier papers4
- Neuron Empirical Gradient: Discovering and Quantifying Neurons' Global Linear ControllabilityXin Zhao, Zehui Jiang, Naoki YoshinagaACL 2025 · 3 citations
- SAEs Are Good for Steering - If You Select the Right FeaturesDana Arad, Aaron Mueller, Yonatan BelinkovEMNLP 2025
- Diff4TST: Masked Diffusion Language Model for Text Style TransferXinchen Ma, Gaole He, Yunshi Lan, Weining QianACL 2026
- Towards a Unified Paradigm of Concept Editing in Large Language ModelsZhuowen Han, Xinwei Wu, Dan Shi, Renren Jin et al.EMNLP 2025
Builds on12
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim et al.ICLR 2024 · 354 citations
- ParaDetox: Detoxification with Parallel DataVarvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy et al.ACL 2022 · 96 citations
- Contrastive Decoding: Open-ended Text Generation as OptimizationXiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang et al.ACL 2023 · 78 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
Related papers
- Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Arbitrary Textual Style Transfer with Small Language ModelsMirac Suzgun, Luke Melas-Kyriazi, Dan JurafskyEMNLP 2022 · 34 citations
- Text Style Transfer Back-TranslationDaimeng Wei, Zhanglin Wu, Hengchao Shang, Zongyao Li et al.ACL 2023 · 10 citations
- Masked Based Unsupervised Content TransferRon Mokady, Sagie Benaim, Lior Wolf, Amit BermanoICLR 2020
- SC2: Towards Enhancing Content Preservation and Style Consistency in Long Text Style TransferJie Zhao, Ziyu Guan, Cai Xu, Wei Zhao et al.ACL 2024
- Rethinking Style Transformer with Energy-based Interpretation: Adversarial Unsupervised Style Transfer using a Pretrained ModelHojun Cho, Dohee Kim, Seungwoo Ryu, ChaeHun Park et al.EMNLP 2022
