Leveraging Outline-Optimized Generative Interactions and Critique for Self-Refining Outlines with Reinforcement Learning
Hengwei Liu, Haoyuan Ma, Qingqing Lyu, Daoxin Zhang, Yao Hu, Yongliang Shen, Yin Zhang, Weiming Lu
Abstract
Long-form outline generation requires satisfying multiple competing objectives simultaneously: outlines must be engaging, wellorganized, topically relevant, and comprehensive while maintaining logical consistency across hierarchical structures. Current approaches either rely on expensive multi-turn interactions with large language models or employ procedural refinement pipelines that cannot systematically learn from critique. We present LOGIC-RL, a framework that transforms critique-guided outline refinement into a learnable policy through reinforcement learning. Our approach constructs refinement trajectories from teacher demonstrations, synthesizes explicit reasoning chains that decompose the critique-revision process, and optimizes a refinement policy using group relative policy optimization with structure-aware rewards. Experiments on FreshWiki and WikiOutline demonstrate that LOGIC-RL achieves substantial improvements over strong baselines, with the 0.6B model obtaining 79.17% relative gain and the 1.7B model achieving 8.67% improvement in average rubric scores compared to the best existing methods. Further analysis reveals that learned refinement policies generalize across domains and can be iteratively applied, with quality continuing to improve through three refinement rounds before diminishing returns.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 81e498d2-c20b-4c2a-9a4e-3cb7722bf756Builds on6
- Prometheus: Inducing Fine-Grained Evaluation Capability in Language ModelsSeungone Kim, Jamin Shin, Yejin Choi, Joel Jang et al.ICLR 2024 · 468 citations
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu et al.EuroSys 2025 · 61 citations
- OmniThink: Expanding Knowledge Boundaries in Machine Writing through ThinkingZekun Xi, Wenbiao Yin, Jizhan Fang, Jialong Wu et al.EMNLP 2025 · 1 citation
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du et al.ICLR 2023
- Generating Biographies on Wikipedia: The Impact of Gender Bias on the Retrieval-Based Generation of Women BiographiesAngela Fan, Claire GardentACL 2022
Related papers
- DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMsYuanhao Li, Mingshan Liu, Hongbo Wang, Yiding Zhang et al.AAAI 2026
- Plan Then Action: High-Level Planning Guidance Reinforcement Learning for LLM ReasoningZhihao Dou, Qinjian Zhao, Zhongwei Wan, Zhang Dinggen et al.ICML 2026 · 24 citations
- MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement LearningZhiheng Xi, Yuhui Wang, Yiwen Ding, Guanyu Li et al.AAAI 2026
- AG-GRPO: Answer-Guided GRPO for Masked Diffusion Language ModelsJuhyeong Kim, Gyunyeop Kim, Sangwoo KangACL 2026
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang et al.NeurIPS 2025 · 387 citations
