Making Sense of Reinforcement Learning and Probabilistic Inference
Brendan O'Donoghue, Ian Osband, Catalin Ionescu
摘要
Reinforcement learning (RL) combines a control problem with statistical estimation: The system dynamics are not known to the agent, but can be learned through experience. A recent line of research casts 'RL as inference' and suggests a particular framework to generalize the RL problem as probabilistic inference. Our paper surfaces a key shortcoming in that approach, and clarifies the sense in which RL can be coherently cast as an inference problem. In particular, an RL agent must consider the effects of its actions upon future rewards and observations: The exploration-exploitation tradeoff. In all but the most simple settings, the resulting inference is computationally intractable so that practical RL algorithms must resort to approximation. We demonstrate that the popular 'RL as inference' approximation can perform poorly in even very basic problems. However, we show that with a small modification the framework does yield algorithms that can provably perform well, and we show that the resulting algorithm is equivalent to the recently proposed K-learning, which we further connect with Thompson sampling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Reward is enough for convex MDPsTom Zahavy, Brendan O'Donoghue, Guillaume Desjardins, Satinder SinghNeurIPS 2021 · 被引用 96 次
- Variational Bayesian Reinforcement Learning with Regret BoundsBrendan O'DonoghueNeurIPS 2021 · 被引用 48 次
- Beyond Verifiable Rewards: Scaling Reinforcement Learning in Language Models to Unverifiable DataYunhao Tang, Sid Wang, Lovish Madaan, Rémi MunosNeurIPS 2025 · 被引用 27 次
- Deciding What to Model: Value-Equivalent Sampling for Reinforcement LearningDilip Arumugam, Benjamin Van RoyNeurIPS 2022 · 被引用 25 次
它引用的顶会 Paper2
相关 Paper
- Probabilistic Inference in Reinforcement Learning Done RightJean Tarbouriech, Tor Lattimore, Brendan O'DonoghueNeurIPS 2023 · 被引用 15 次
- Q-learning with Posterior SamplingPriyank Agrawal, Shipra Agrawal, Azmat AzatiICLR 2026 · 被引用 3 次
- Provably Efficient Exploration in Inverse Constrained Reinforcement LearningBo Yue, Jian Li, Guiliang LiuICML 2025
- Making RL with Preference-based Feedback Efficient via RandomizationRunzhe Wu, Wen SunICLR 2024 · 被引用 44 次
- Outcome-Driven Reinforcement Learning via Variational InferenceTim G. J. Rudner, Vitchyr Pong, Rowan McAllister, Yarin Gal 等NeurIPS 2021 · 被引用 24 次
