Dynamic Multi-Reward Weighting for Multi-Style Controllable Generation
Karin de Langis, Ryan Koo, Dongyeop Kang
Abstract
Textual style expresses a diverse set of information, including interpersonal dynamics (e.g., formality) and the author's emotions or attitudes (e.g., disgust ). An open question is how language models can be explicitly controlled so that they weave together target styles when generating text: for example, to produce text that is both negative and non-toxic. One approach to such controlled generation is multiobjective reinforcement learning (RL), but how to best combine multiple objectives in a reward function is an open question. In this paper, we investigate various formulations of multi-style rewards, including calibrated outputs from discriminators and dynamic weighting by discriminator gradient magnitudes. We find that our proposed dynamic weighting outperforms static weighting approaches with respect style control while maintaining linguistic quality, and we explore its effectiveness in 2-and 3-style control. All code and data for the RL pipelines will be publicly available. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 086ee03c-d160-46fd-a0eb-1b623bfd725bCited by top-tier papers3
- LLM-based Rewriting of Inappropriate Argumentation using Reinforcement Learning from Machine FeedbackTimon Ziegenbein, Gabriella Skitalinskaya, Alireza Bayat Makou, Henning WachsmuthACL 2024
- Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-TuningYanbei Jiang, Amr Keleg, Ryandito Diandaru, Jey Han Lau et al.ACL 2026
- Pareto-Guided Optimal Transport for Multi-Reward AlignmentYing Ba, Tianyu Zhang, Mohan Zhou, Yalong Bai et al.ICML 2026
Builds on9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri et al.NeurIPS 2023 · 516 citations
Related papers
- Preference-Controlled Multi-Objective Reinforcement Learning for Conditional Text GenerationWenqing Chen, Jidong Tian, Caoyun Fan, Yitian Li et al.AAAI 2023 · 2 citations
- DORB: Dynamically Optimizing Multiple Rewards with BanditsRamakanth Pasunuru, Han Guo, Mohit BansalEMNLP 2020 · 2 citations
- Controlled Text Generation as Continuous Optimization with Multiple ConstraintsSachin Kumar, Eric Malmi, Aliaksei Severyn, Yulia TsvetkovNeurIPS 2021 · 91 citations
- Diff4TST: Masked Diffusion Language Model for Text Style TransferXinchen Ma, Gaole He, Yunshi Lan, Weining QianACL 2026
- Controlled Decoding from Language ModelsSidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li et al.ICML 2024 · 130 citations
