Generalizing Consistency Policy to Visual RL with Prioritized Proximal Experience Regularization
Haoran Li, Zhennan Jiang, Yuhui Chen, Dongbin Zhao
Abstract
With high-dimensional state spaces, visual reinforcement learning (RL) faces significant challenges in exploitation and exploration, resulting in low sample efficiency and training stability. As a time-efficient diffusion model, although consistency models have been validated in online state-based RL, it is still an open question whether it can be extended to visual RL. In this paper, we investigate the impact of non-stationary distribution and the actor-critic framework on consistency policy in online RL, and find that consistency policy was unstable during the training, especially in visual RL with the high-dimensional state space. To this end, we suggest sample-based entropy regularization to stabilize the policy training, and propose a consistency policy with prioritized proximal experience regularization (CP3ER) to improve sample efficiency. CP3ER achieves new state-of-the-art (SOTA) performance in 21 tasks across DeepMind control suite and Meta-world. To our knowledge, CP3ER is the first method to apply diffusion/consistency models to visual RL and demonstrates the potential of consistency models in visual RL. More visualization results are available at https://jzndd.github.io/CP3ER-Page/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1674ccee-e617-4687-907c-28761b9dba50Cited by top-tier papers6
- MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous DrivingJunli Wang, Yinan Zheng, Xueyi Liu, Zebin Xing et al.CVPR 2026 · 16 citations
- Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory DistillationXintong Duan, Yutong He, Fahim Tajwar, Russ Salakhutdinov et al.ICLR 2026 · 3 citations
- Task-Aware Exploration via a Predictive Bisimulation MetricDayang Liang, Ruihan LIU, Lipeng Wan, Yunlong Liu et al.ICML 2026 · 1 citation
- Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement LearningJingbo Sun, Qichao Zhang, Songjun Tu, Xing Fang et al.CVPR 2026 · 1 citation
- Unsupervised Zero-Shot Reinforcement Learning via Dual-Value Forward-Backward RepresentationJingbo Sun, Songjun Tu, Qichao Zhang, Haoran Li et al.ICLR 2025
Builds on31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
Related papers
- Value-Consistent Representation Learning for Data-Efficient Reinforcement LearningYang Yue, Bingyi Kang, Zhongwen Xu, Gao Huang et al.AAAI 2023 · 19 citations
- Consistency Models as a Rich and Efficient Policy Class for Reinforcement LearningZihan Ding, Chi JinICLR 2024 · 73 citations
- What Effects the Generalization in Visual Reinforcement Learning: Policy Consistency with Truncated Return PredictionShuo Wang, Zhihao Wu, Xiaobo Hu, Jinwen Wang et al.AAAI 2024 · 18 citations
- When Should We Prefer State-to-Visual DAgger over Visual Reinforcement Learning?Tongzhou Mu, Zhaoyang Li, Stanislaw Wiktor Strzelecki, Xiu Yuan et al.AAAI 2025 · 8 citations
- Simplified Temporal Consistency Reinforcement LearningYi Zhao, Wenshuai Zhao, Rinu Boney, Juho Kannala et al.ICML 2023 · 19 citations
