Group Causal Policy Optimization for Post-Training Large Language Models
Ziyin Gu, Jingyao Wang, Ran Zuo, Chuxiong Sun, Zeen Song, Changwen Zheng, Wenwen Qiang
摘要
Recent advances in large language models (LLMs) have broadened their applicability across diverse tasks, yet specialized domains still require targeted post-training. Among existing methods, Group Relative Policy Optimization (GRPO) stands out for its efficiency, leveraging groupwise relative rewards while avoiding costly value function learning. However, GRPO treats candidate responses as independent, overlooking semantic interactions such as complementarity and contradiction. To address this challenge, we first introduce a Structural Causal Model (SCM) that reveals hidden dependencies among candidate responses induced by conditioning on a final integrated output-forming a collider structure. Then, our causal analysis leads to two insights: (1) projecting responses onto a causally-informed subspace improves prediction quality, and (2) this projection yields a better baseline than query-only conditioning. Building on these insights, we propose Group Causal Policy Optimization (GCPO), which integrates causal structure into optimization through two key components: a causally-informed reward adjustment and a novel KL-regularization term that aligns the policy with a causally-projected reference distribution. Comprehensive experimental evaluations demonstrate that GCPO consistently surpasses existing methods-including GRPO-across multiple reasoning benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- On the Plasticity and Stability for Post-Training Large Language ModelsWenwen Qiang, Ziyin Gu, Jiahuan Zhou, Jie Hu 等ICML 2026 · 被引用 3 次
- COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMsPeizheng Guo, Jingyao Wang, Wenwen Qiang, Jiahuan Zhou 等CVPR 2026 · 被引用 1 次
- Pareto-Guided Optimal Transport for Multi-Reward AlignmentYing Ba, Tianyu Zhang, Mohan Zhou, Yalong Bai 等ICML 2026
它引用的顶会 Paper8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree SearchDan Zhang, Sining Zhoubian, Ziniu Hu, Yisong Yue 等NeurIPS 2024 · 被引用 527 次
- DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement LearningHao Bai, Yifei Zhou, Jiayi Pan, Mert Cemri 等NeurIPS 2024 · 被引用 239 次
- GVPO: Group Variance Policy Optimization for Large Language Model Post-TrainingKaichen Zhang, Yuzhong Hong, Junwei Bao, Hongfei Jiang 等NeurIPS 2025 · 被引用 35 次
相关 Paper
- Group-Aware Reinforcement Learning for Output Diversity in Large Language ModelsOron Anschel, Alon Shoshan, Adam Botach, Shunit Haviv Hakimi 等EMNLP 2025 · 被引用 1 次
- Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy OptimizationYifeng Ding, Hung Le, Songyang Han, Kangrui Ruan 等ACL 2026 · 被引用 5 次
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL OptimizationShih-Yang Liu, Xin Dong, Ximing Lu, Shizhe Diao 等ICML 2026 · 被引用 128 次
- RiskPO: Risk-based Policy Optimization with Verifiable Reward for LLM Post-TrainingTao Ren, Jinyang Jiang, Hui Yang, Wan Tian 等ICLR 2026 · 被引用 8 次
- MVP: Enhancing Video Large Language Models via Self-supervised Masked Video PredictionXiaokun Sun, Zezhong Wu, Zewen Ding, Linli XuACL 2026 · 被引用 1 次
