GIPO: Gaussian Importance Sampling Policy Optimization
程烜 陆, Zhenquan Zhang, Shukuan Wang, Qunzhi Lin, Baigui Sun, Yang Liu
Abstract
Post-training with reinforcement learning (RL) has recently shown strong promise for advancing multimodal agents beyond supervised imitation. However, RL remains limited by poor data efficiency, particularly in settings where interaction data are scarce and quickly become outdated. To address this challenge, GIPO (Gaussian Importance sampling Policy Optimization) is proposed as a policy optimization objective based on truncated importance sampling, replacing hard clipping with a log-ratio-based Gaussian trust weight to softly damp extreme importance ratios while maintaining non-zero gradients. Theoretical analysis shows that GIPO introduces an implicit, tunable constraint on the update magnitude, while concentration bounds guarantee robustness and stability under finite-sample estimation. Experimental results show that GIPO achieves state-of-the-art performance among clipping-based baselines across a wide range of replay buffer sizes, from near on-policy to highly stale data, while exhibiting superior bias--variance trade-off, high training stability and improved sample efficiency. Code is available at https://github.com/distanceLu/GIPO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2467b5de-4cf4-4b26-9695-0d1db902e217Builds on2
Related papers
- Ratio-Variance Regularized Policy OptimizationYu Luo, Shuo Han, Yihan Hu, Lei Lv et al.ICML 2026
- GVPO: Group Variance Policy Optimization for Large Language Model Post-TrainingKaichen Zhang, Yuzhong Hong, Junwei Bao, Hongfei Jiang et al.NeurIPS 2025 · 35 citations
- Truncating Trajectories in Monte Carlo Reinforcement LearningRiccardo Poiani, Alberto Maria Metelli, Marcello RestelliICML 2023 · 6 citations
- Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy TrainingYoussef Mroueh, Nicolas Dupuis, Brian Belgodere, Apoorva Nitsure et al.ICLR 2026 · 39 citations
- MARPO: A Reflective Policy Optimization for Multi-Agent Reinforcement LearningCuiling Wu, Yaozhong Gan, Junliang Xing, Ying FuAAAI 2026
