Boosting Multi-Domain Reasoning of LLMs via Curvature-Guided Policy Optimization
Xize Liang, Lin Yang, Jie Wang, Rui Liu, Yang Lu, Jinliang Zeng, Hanzhu Chen, Dong Li, Jianye HAO
摘要
Multi-domain reinforcement learning (RL) for large language models (LLMs) involves highly intricate reward surfaces, posing significant challenges in finding parameters that excel across all domains. Recent empirical studies have further highlighted conflicts among domains, where gains in one capability often come at the expense of another. However, approaches to mitigate such conflicts and enhance multi-domain reasoning remain largely underexplored. To address this challenge, we propose Curvature-Guided Policy Optimization (CGPO), a principled and scalable training framework to advance the multi-domain reasoning of LLMs. Inspired by Newton's method, CGPO exploits the geometric structure in the reward surface, while sidestepping the prohibitive cost of Hessian computation. At each update, CGPO processes domains in random order, preconditioning their gradients with curvature information from other domains to foster richer cross-domain interactions. This mechanism further promotes implicit gradient alignment by maximizing inter-domain inner products in expectation, steering the parameters toward regions that jointly enhance multi-domain performance. Extensive experiments on a mixed dataset covering math, coding, science, and creative writing, evaluated across seven widely-used benchmarks, show that CGPO significantly outperforms all baselines in terms of faster reward improvement and stronger multi-domain capability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable RewardsYiran Shen, Yu Xia, Jonathan Chang, Prithviraj AmmanabroluICML 2026
- Evolving Graph Structured Programs for Circuit Generation with Large Language ModelsYinqi Bai, Jie Wang, Lei Chen, Zhihai Wang 等ICLR 2026
- D-ARL: A Distribution-Matched Asynchronous Reinforcement Learning Framework for Language Reasoning白 寅岐, Xialiang Tong, Jie Wang, Hongyu Liu 等ICML 2026
它引用的顶会 Paper21
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone 等NeurIPS 2021 · 被引用 686 次
相关 Paper
- Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain ReasoningBaolong Bi, Shenghua Liu, Yiwei Wang, Siqian Tong 等ICML 2026 · 被引用 19 次
- Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM ReasoningLuckeciano Carvalho Melo, Alessandro Abate, Yarin GalICLR 2026 · 被引用 7 次
- Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy OptimizationYifeng Ding, Hung Le, Songyang Han, Kangrui Ruan 等ACL 2026 · 被引用 5 次
- From Imitation to Discrimination: Toward a Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning TasksChangpeng Yang, Jinyang Wu, Yuchen Liu, Shuai Zhang 等AAAI 2026
- Stronger-MAS: Multi-Agent Reinforcement Learning for Collaborative LLMsYujie Zhao, Lanxiang Hu, Yang Wang, Minmin Hou 等ICLR 2026 · 被引用 26 次
