Dynamic Multi-Reward Weighting for Multi-Style Controllable Generation
Karin de Langis, Ryan Koo, Dongyeop Kang
摘要
Textual style expresses a diverse set of information, including interpersonal dynamics (e.g., formality) and the author's emotions or attitudes (e.g., disgust ). An open question is how language models can be explicitly controlled so that they weave together target styles when generating text: for example, to produce text that is both negative and non-toxic. One approach to such controlled generation is multiobjective reinforcement learning (RL), but how to best combine multiple objectives in a reward function is an open question. In this paper, we investigate various formulations of multi-style rewards, including calibrated outputs from discriminators and dynamic weighting by discriminator gradient magnitudes. We find that our proposed dynamic weighting outperforms static weighting approaches with respect style control while maintaining linguistic quality, and we explore its effectiveness in 2-and 3-style control. All code and data for the RL pipelines will be publicly available. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LLM-based Rewriting of Inappropriate Argumentation using Reinforcement Learning from Machine FeedbackTimon Ziegenbein, Gabriella Skitalinskaya, Alireza Bayat Makou, Henning WachsmuthACL 2024
- Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-TuningYanbei Jiang, Amr Keleg, Ryandito Diandaru, Jey Han Lau 等ACL 2026
- Pareto-Guided Optimal Transport for Multi-Reward AlignmentYing Ba, Tianyu Zhang, Mohan Zhou, Yalong Bai 等ICML 2026
它引用的顶会 Paper9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri 等NeurIPS 2023 · 被引用 516 次
相关 Paper
- Preference-Controlled Multi-Objective Reinforcement Learning for Conditional Text GenerationWenqing Chen, Jidong Tian, Caoyun Fan, Yitian Li 等AAAI 2023 · 被引用 2 次
- DORB: Dynamically Optimizing Multiple Rewards with BanditsRamakanth Pasunuru, Han Guo, Mohit BansalEMNLP 2020 · 被引用 2 次
- Controlled Text Generation as Continuous Optimization with Multiple ConstraintsSachin Kumar, Eric Malmi, Aliaksei Severyn, Yulia TsvetkovNeurIPS 2021 · 被引用 91 次
- Diff4TST: Masked Diffusion Language Model for Text Style TransferXinchen Ma, Gaole He, Yunshi Lan, Weining QianACL 2026
- Controlled Decoding from Language ModelsSidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li 等ICML 2024 · 被引用 130 次
