DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative Writing
Qian Cao, Yahui Liu, Wei Bi, Yi Zhao, Ruihua Song, Xiting Wang, Ruiming Tang, Guorui Zhou, Han Li
Abstract
Reinforcement learning (RL)-based enhancement of large language models (LLMs) often leads to reduced output diversity, undermining their utility in open-ended tasks like creative writing. Current methods lack explicit mechanisms for guiding diverse exploration and instead prioritize optimization efficiency and performance over diversity. This paper proposes an RL framework structured around a semi-structured long Chain-of-Thought (CoT), in which the generation process is decomposed into explicitly planned intermediate steps. We introduce a Diverse Planning Branching method that strategically introduces divergence at the planning phase based on diversity variation, alongside a group-aware diversity reward to encourage distinct trajectories. Experimental results on creative writing benchmarks demonstrate that our approach significantly improves output diversity without compromising generation quality, consistently outperforming existing baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Understanding the Effects of RLHF on LLM Generalisation and DiversityRobert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina et al.ICLR 2024 · 332 citations
- Does Writing with Language Models Reduce Content Diversity?Vishakh Padmakumar, He HeICLR 2024 · 173 citations
- Contrastive Decoding: Open-ended Text Generation as OptimizationXiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang et al.ACL 2023 · 78 citations
Related papers
- IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural ThinkingZechen Sun, Yuyang Sun, Zecheng Tang, Juntao Li et al.ACL 2026
- DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMsYuanhao Li, Mingshan Liu, Hongbo Wang, Yiding Zhang et al.AAAI 2026
- Curiosity-Driven Reinforcement Learning from Human FeedbackHaoran Sun, Yekun Chai, Shuohuan Wang, Yu Sun et al.ACL 2025 · 17 citations
- Automatic Chain of Thought Prompting in Large Language ModelsZhuosheng Zhang, Aston Zhang, Mu Li, Alex SmolaICLR 2023 · 234 citations
- RLMR: Reinforcement Learning with Mixed Rewards for Creative WritingJianxing Liao, Tian Zhang, Xiao Feng, Yusong Zhang et al.AAAI 2026 · 6 citations
