StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of Style Elements
Jillian Fisher, Skyler Hallinan, Ximing Lu, Mitchell L. Gordon, Zaïd Harchaoui, Yejin Choi
Abstract
Authorship obfuscation, rewriting a text to intentionally obscure the identity of the author, is an important but challenging task. Current methods using large language models (LLMs) lack interpretability and controllability, often ignoring author-specific stylistic features, resulting in less robust performance overall. To address this, we develop STYLEREMIX, an adaptive and interpretable obfuscation method that perturbs specific, fine-grained style elements of the original input text. STYLEREMIX uses pre-trained Low Rank Adaptation (LoRA) modules to rewrite an input specifically along various stylistic axes (e.g., formality and length) while maintaining low computational cost. STYLEREMIX outperforms state-of-theart baselines and much larger LLMs in a variety of domains as assessed by both automatic and human evaluation. Additionally, we release AUTHORMIX, a large set of 30K high-quality, long-form texts from a diverse set of 14 authors and 4 domains, and DISC, a parallel corpus of 1,500 texts spanning seven style axes in 16 unique directions 1 . * Co-first authors 1 We release 1) our code at https://github.com/ jfisher52/StyleRemix 2) a demo of STYLEREMIX at https://huggingface.co/spaces/hallisky/ StyleRemix and 3) the datasets (AUTHORMIX and DISC) and trained models in a HuggingFace collection
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Personalized Text Generation with Contrastive Activation SteeringJinghao Zhang, Yuting Liu, Wenjie Wang, Qiang Liu et al.ACL 2025 · 24 citations
- Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and DomainsJunghwan Kim, Haotian Zhang, David JurgensEMNLP 2025 · 3 citations
- FedMerge: Federated Model Merging for PersonalizationShutong Chen, Tianyi Zhou, Guodong Long, Jing Jiang et al.AAAI 2026 · 2 citations
- Unraveling Interwoven Roles of Large Language Models in Authorship Privacy: Obfuscation, Mimicking, and VerificationTuc Nguyen, Yifan Hu, Thai LeEMNLP 2025
Builds on14
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LunchLe Yu, Bowen Yu, Haiyang Yu, Fei Huang et al.ICML 2024 · 605 citations
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
Related papers
- A Girl Has A Name: Detecting Authorship ObfuscationAsad Mahmood, Zubair Shafiq, Padmini SrinivasanACL 2020 · 1 citation
- Causal-Steer: Disentangled Continuous Style Control without Parallel CorporaQingsong Wang, Chang Yao, Jingyuan ChenICLR 2026
- ALISON: Fast and Effective Stylometric Authorship ObfuscationEric Xing, Saranya Venkatraman, Thai Le, Dongwon LeeAAAI 2024 · 7 citations
- Adversarial Authorship Attribution for DeobfuscationWanyue Zhai, Jonathan Rusert, Zubair Shafiq, Padmini SrinivasanACL 2022 · 7 citations
- Adapting Language Models for Non-Parallel Author-Stylized RewritingBakhtiyar Syed, Gaurav Verma, Balaji Vasan Srinivasan, Anandhavelu Natarajan et al.AAAI 2020 · 53 citations
