Policy Gradient With Serial Markov Chain Reasoning
Edoardo Cetin, Oya Çeliktutan
摘要
We introduce a new framework that performs decision-making in reinforcement learning (RL) as an iterative reasoning process. We model agent behavior as the steady-state distribution of a parameterized reasoning Markov chain (RMC), optimized with a new tractable estimate of the policy gradient. We perform action selection by simulating the RMC for enough reasoning steps to approach its steady-state distribution. We show our framework has several useful properties that are inherently missing from traditional RL. For instance, it allows agent behavior to approximate any continuous distribution over actions by parameterizing the RMC with a simple Gaussian transition function. Moreover, the number of reasoning steps to reach convergence can scale adaptively with the difficulty of each action selection decision and can be accelerated by re-using past solutions. Our resulting algorithm achieves state-of-the-art performance in popular Mujoco and DeepMind Control benchmarks, both for proprioceptive and pixel-based tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- S2AC: Energy-Based Reinforcement Learning with Stein Soft Actor CriticSafa Messaoud, Billel Mokeddem, Zhenghai Xue, Linsey Pang 等ICLR 2024 · 被引用 21 次
- Learning Intractable Multimodal Policies with Reparameterization and Diversity RegularizationZiqi Wang, Jiashun Liu, Ling PanNeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper10
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 被引用 911 次
相关 Paper
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez 等ICLR 2021 · 被引用 77 次
- Stop! Planner Time: Metareasoning for Probabilistic Planning Using Learned Performance ProfilesMatthew Budd, Bruno Lacerda, Nick HawesAAAI 2024 · 被引用 2 次
- Diffusion Actor-Critic with Entropy RegulatorYinuo Wang, Likun Wang, Yuxuan Jiang, Wenjun Zou 等NeurIPS 2024 · 被引用 105 次
- Recursive Reasoning Graph for Multi-Agent Reinforcement LearningXiaobai Ma, David Isele, Jayesh K. Gupta, Kikuo Fujimura 等AAAI 2022 · 被引用 8 次
- Distributions as Actions: A Unified Framework for Diverse Action SpacesJiamin He, A. Rupam Mahmood, Martha WhiteICLR 2026
