Learning from Comparison: Constrained Projection Policy Optimization for Pareto-Front Improvement
Jintao Li, Maowen Tang, Yongji Long, Weixuan Liu, Yanlang Zheng, Sicheng He, Ao-Jin Li, Shui Yu, Yun Li
摘要
Constrained multi-objective reinforcement learning aims to discover a diverse set of feasible trade-offs, yet scalarization and signed, normalized group-relative advantages can be brittle under objective-scale drift, near-ties, and feasibility scarcity. We propose constrained projection policy optimization (CoPro), which alternates between an E-step moment projection and an M-step policy projection. In the E-step, we solve a Kullback-Leibler (KL)-regularized, moment-constrained projection over each sampled group to compute a nonnegative reweighting distribution () that promotes feasible Pareto-front (PF) progress, preserves feasibility anchors, and suppresses ambiguous near-ties. This E-step admits a closed-form exponential-family solution and guarantees strictly positive probability mass on feasible anchors whenever feasible candidates appear in the group. In the M-step, we project the policy toward via weighted maximum likelihood with a trust-region regularizer, yielding a PF-aligned update direction from comparisons without hand-crafted reward shaping. Empirically, CoPro improves feasible PF quality and robustness on constrained multi-objective benchmarks for large language model tool use and analog circuit design tasks; code is available at CoPro.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 被引用 403 次
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang 等NeurIPS 2025 · 被引用 387 次
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus 等ICML 2020 · 被引用 210 次
相关 Paper
- BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement LearningYuan Li, Bo Wang, Yufei Gao, Yuqian Yao 等ICML 2026 · 被引用 2 次
- MRPO: Magnitude-Regularized Policy Optimization via L1 ConstraintsWei Han, Yuanxing Liu, Mingda Li, Ruiyu Xiao 等ICML 2026
- MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction FollowingMohammad Mahdi Salmani-Zarchi, Zahra Rahimi, Heshaam Faili, Mohammad Javad DoustiACL 2026
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 被引用 306 次
- Proactive Constrained Policy Optimization with Preemptive PenaltyNing Yang, Pengyu Wang, Guoqing Liu, Haifeng Zhang 等AAAI 2026 · 被引用 1 次
