RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing
Jianxing Liao, Tian Zhang, Xiao Feng, Yusong Zhang, Haorui Wang, Bosi Wen, Ziying Wang, Runzhi Shi
摘要
Large language models are extensively utilized in creative writing applications. Creative writing requires a balance between subjective writing quality (e.g., literariness and emotional expression) and objective constraint following (e.g., format requirements and word limits). Existing reinforcement learning methods struggle to balance these two aspects: single reward strategies fail to improve both abilities simultaneously, while fixed-weight mixed-reward methods lack the ability to adapt to different writing scenarios. To address this problem, we propose Reinforcement Learning with Mixed Rewards (RLMR), utilizing a dynamically mixed reward system from a writing reward model evaluating subjective writing quality and a constraint verification model assessing objective constraint following. The constraint following reward weight is adjusted dynamically according to the writing quality within sampled groups, ensuring that samples violating constraints get negative advantage in GRPO and thus penalized during training, which is the key innovation of this proposed method. We conduct automated and manual evaluations across diverse model families from 8B to 72B parameters. Additionally, we construct a real-world writing benchmark named WriteEval for comprehensive evaluation. Results illustrate that our method achieves consistent improvements in both instruction following (IFEval from 83.36% to 86.65%) and writing quality (72.75% win rate in manual expert pairwise evaluations on WriteEval). To the best of our knowledge, RLMR is the first work to combine subjective preferences with objective verification in online RL training, providing an effective solution for multi-dimensional creative writing optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- WildReward: Learning Reward Models from In-the-Wild Human InteractionsHao Peng, Yunjia Qi, Xiaozhi Wang, Zijun Yao 等ACL 2026 · 被引用 3 次
- SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data UtilityXuyang Zhi, Peilun Zhou, Chengqiang Lu, Hang Lv 等ACL 2026 · 被引用 3 次
- Rubric Curriculum RL: Exploiting the Generation-Verification Gap in Non-Verifiable DomainsTejas Krishnan, Sumeet Motwani, Charles London, Suhaas Bhat 等ICML 2026
它引用的顶会 Paper6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler 等NeurIPS 2020 · 被引用 124 次
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu 等EuroSys 2025 · 被引用 61 次
- LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement LearningYuhao Wu, Yushi Bai, Zhiqiang Hu, Roy Ka-Wei Lee 等ICLR 2026 · 被引用 9 次
相关 Paper
- Confronting Reward Model Overoptimization with Constrained RLHFTed Moskovitz, Aaditya K. Singh, DJ Strouse, Tuomas Sandholm 等ICLR 2024 · 被引用 89 次
- Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement LearningXuanyu Lei, Chenliang Li, Yuning Wu, Kaiming Liu 等ACL 2026 · 被引用 8 次
- Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined RewardsXiaolong Wei, Bo Lu, Xingyu Zhang, Zhejun Zhao 等EMNLP 2025
- Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-TrainingRan Xu, Tianci Liu, Zihan Dong, Tony Yu 等ICML 2026
- Teaching Models to Improve on TapeLiat Bezalel, Eyal Orgad, Amir GlobersonAAAI 2025
