LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
Yuhao Wu, Yushi Bai, Zhiqiang Hu, Roy Ka-Wei Lee, Juanzi Li
Abstract
Ultra-long generation by large language models (LLMs) is a widely demanded scenario, yet it remains a significant challenge due to their maximum generation length limit and overall quality degradation as sequence length increases. Previous approaches, exemplified by LongWriter, typically rely on ''teaching'', which involves supervised fine-tuning (SFT) on synthetic long-form outputs. However, this strategy heavily depends on synthetic SFT data, which is difficult and costly to construct, often lacks coherence and consistency, and tends to be overly artificial and structurally monotonous. In this work, we propose an incentivization-based approach that, starting entirely from scratch and without relying on any annotated or synthetic data, leverages reinforcement learning (RL) to foster the emergence of ultra-long, high-quality text generation capabilities in LLMs. We perform RL training starting from a base model, similar to R1-Zero, guiding it to engage in reasoning that facilitates planning and refinement during the writing process. To support this, we employ specialized reward models that steer the LLM towards improved length control, writing quality, and structural formatting. Experimental evaluations show that our LongWriter-Zero model, trained from Qwen2.5-32B, consistently outperforms traditional SFT methods on long-form writing tasks, achieving state-of-the-art results across all metrics on WritingBench and Arena-Write, and even surpassing 100B+ models such as DeepSeek R1 and Qwen3-235B.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 472072d2-e691-47bd-b74a-9fe5ce79eb24Cited by top-tier papers6
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for Open-Ended LLM ReasoningYang Zhou, Sunzhu Li, Shunyu Liu, Wenkai Fang et al.ICML 2026 · 44 citations
- FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft LearningYizhou Zhang, Ning Lv, Teng Wang, Jisheng DangICLR 2026 · 10 citations
- Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement LearningXuanyu Lei, Chenliang Li, Yuning Wu, Kaiming Liu et al.ACL 2026 · 8 citations
- RLMR: Reinforcement Learning with Mixed Rewards for Creative WritingJianxing Liao, Tian Zhang, Xiao Feng, Yusong Zhang et al.AAAI 2026 · 6 citations
- Long-form RewardBench: Evaluating Reward Models for Long-form GenerationHui Huang, Yancheng He, Wei Liu, Muyun Yang et al.AAAI 2026 · 1 citation
Builds on13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang et al.NeurIPS 2025 · 1,109 citations
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie et al.ICLR 2024 · 504 citations
Related papers
- LongWriter: Unleashing 10, 000+ Word Generation from Long Context LLMsYushi Bai, Jiajie Zhang, Xin Lv, Linzhi Zheng et al.ICLR 2025
- NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement LearningWei Liu, Siya Qi, Xinyu Wang, Chen Qian et al.EMNLP 2025 · 4 citations
- LongWriter-V: Enabling Ultra-Long and High-Fidelity Generation in Vision-Language ModelsShangqing Tu, Yucheng Wang, Daniel Zhang-Li, Yushi Bai et al.ACM MM 2025
- General-Reasoner: Advancing LLM Reasoning Across All DomainsXueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang et al.NeurIPS 2025 · 153 citations
- DeepWriter: A Multi-Agent Collaboration Framework for Information-rich Ultra-long Book WritingMing Wang, Minghao Hu, Xiuli Kang, Li He et al.AAAI 2026
