Generalizing Consistency Policy to Visual RL with Prioritized Proximal Experience Regularization
Haoran Li, Zhennan Jiang, Yuhui Chen, Dongbin Zhao
摘要
With high-dimensional state spaces, visual reinforcement learning (RL) faces significant challenges in exploitation and exploration, resulting in low sample efficiency and training stability. As a time-efficient diffusion model, although consistency models have been validated in online state-based RL, it is still an open question whether it can be extended to visual RL. In this paper, we investigate the impact of non-stationary distribution and the actor-critic framework on consistency policy in online RL, and find that consistency policy was unstable during the training, especially in visual RL with the high-dimensional state space. To this end, we suggest sample-based entropy regularization to stabilize the policy training, and propose a consistency policy with prioritized proximal experience regularization (CP3ER) to improve sample efficiency. CP3ER achieves new state-of-the-art (SOTA) performance in 21 tasks across DeepMind control suite and Meta-world. To our knowledge, CP3ER is the first method to apply diffusion/consistency models to visual RL and demonstrates the potential of consistency models in visual RL. More visualization results are available at https://jzndd.github.io/CP3ER-Page/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous DrivingJunli Wang, Yinan Zheng, Xueyi Liu, Zebin Xing 等CVPR 2026 · 被引用 16 次
- Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory DistillationXintong Duan, Yutong He, Fahim Tajwar, Russ Salakhutdinov 等ICLR 2026 · 被引用 3 次
- Task-Aware Exploration via a Predictive Bisimulation MetricDayang Liang, Ruihan LIU, Lipeng Wan, Yunlong Liu 等ICML 2026 · 被引用 1 次
- Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement LearningJingbo Sun, Qichao Zhang, Songjun Tu, Xing Fang 等CVPR 2026 · 被引用 1 次
- Unsupervised Zero-Shot Reinforcement Learning via Dual-Value Forward-Backward RepresentationJingbo Sun, Songjun Tu, Qichao Zhang, Haoran Li 等ICLR 2025
它引用的顶会 Paper31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 被引用 1,720 次
相关 Paper
- Value-Consistent Representation Learning for Data-Efficient Reinforcement LearningYang Yue, Bingyi Kang, Zhongwen Xu, Gao Huang 等AAAI 2023 · 被引用 19 次
- Consistency Models as a Rich and Efficient Policy Class for Reinforcement LearningZihan Ding, Chi JinICLR 2024 · 被引用 73 次
- What Effects the Generalization in Visual Reinforcement Learning: Policy Consistency with Truncated Return PredictionShuo Wang, Zhihao Wu, Xiaobo Hu, Jinwen Wang 等AAAI 2024 · 被引用 18 次
- When Should We Prefer State-to-Visual DAgger over Visual Reinforcement Learning?Tongzhou Mu, Zhaoyang Li, Stanislaw Wiktor Strzelecki, Xiu Yuan 等AAAI 2025 · 被引用 8 次
- Simplified Temporal Consistency Reinforcement LearningYi Zhao, Wenshuai Zhao, Rinu Boney, Juho Kannala 等ICML 2023 · 被引用 19 次
