LLM-based Rewriting of Inappropriate Argumentation using Reinforcement Learning from Machine Feedback
Timon Ziegenbein, Gabriella Skitalinskaya, Alireza Bayat Makou, Henning Wachsmuth
Abstract
Ensuring that online discussions are civil and productive is a major challenge for social media platforms. Such platforms usually rely both on users and on automated detection tools to flag inappropriate arguments of other users, which moderators then review. However, this kind of post-hoc moderation is expensive and time-consuming, and moderators are often overwhelmed by the amount and severity of flagged content. Instead, a promising alternative is to prevent negative behavior during content creation. This paper studies how inappropriate language in arguments can be computationally mitigated. We propose a reinforcement learningbased rewriting approach that balances content preservation and appropriateness based on existing classifiers, prompting an instructionfinetuned large language model (LLM) as our initial policy. Unlike related style transfer tasks, rewriting inappropriate arguments allows deleting and adding content permanently. It is therefore tackled on document level rather than sentence level. We evaluate different weighting schemes for the reward function in both absolute and relative human assessment studies. Systematic experiments on non-parallel data provide evidence that our approach can mitigate the inappropriateness of arguments while largely preserving their content. It significantly outperforms competitive baselines, including few-shot learning, prompting, and humans.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14eb7a5f-2c7c-4d34-84ba-3d4eaaee8a2fCited by top-tier papers3
- Sustaining Human Agency, Attending to Its Cost: An Investigation into Generative AI Design for Non-Native Speakers' Language UseYimin Xiao, Cartor Hancock, Sweta Agrawal, Nikita Mehandru et al.CHI 2025 · 18 citations
- AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive ContextsEsra Dönmez, Maximilian Maurer, Gabriella Lapesa, Agnieszka FalenskaEMNLP 2025 · 1 citation
- Teaching LLMs Human-Like Editing of Inappropriate Argumentation via Reinforcement LearningTimon Ziegenbein, Maja Stahl, Henning WachsmuthACL 2026
Builds on13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler et al.NeurIPS 2020 · 124 citations
Related papers
- Outcome-Constrained Large Language Models for Countering Hate SpeechLingzi Hong, Pengcheng Luo, Eduardo Blanco, Xiaoying SongEMNLP 2024 · 5 citations
- RLPrompt: Optimizing Discrete Text Prompts with Reinforcement LearningMingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang et al.EMNLP 2022 · 141 citations
- Risk-Averse Fine-tuning of Large Language ModelsSapana Chaudhary, Ujwal Dinesha, Dileep Kalathil, Srinivas ShakkottaiNeurIPS 2024 · 12 citations
- Learning to Rewrite Prompts for Personalized Text GenerationCheng Li, Mingyang Zhang, Qiaozhu Mei, Weize Kong et al.WWW 2024 · 54 citations
- Unveiling the Implicit Toxicity in Large Language ModelsJiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang et al.EMNLP 2023 · 21 citations
