Sequential Transfer in Reinforcement Learning with a Generative Model
Andrea Tirinzoni, Riccardo Poiani, Marcello Restelli
摘要
We are interested in how to design reinforcement learning agents that provably reduce the sample complexity for learning new tasks by transferring knowledge from previously-solved ones. The availability of solutions to related problems poses a fundamental trade-off: whether to seek policies that are expected to achieve high (yet suboptimal) performance in the new task immediately or whether to seek information to quickly identify an optimal solution, potentially at the cost of poor initial behavior. In this work, we focus on the second objective when the agent has access to a generative model of state-action pairs. First, given a set of solved tasks containing an approximation of the target one, we design an algorithm that quickly identifies an accurate solution by seeking the state-action pairs that are most informative for this purpose. We derive PAC bounds on its sample complexity which clearly demonstrate the benefits of using this kind of prior knowledge. Then, we show how to learn these approximate tasks sequentially by reducing our transfer setting to a hidden Markov model and employing spectral methods to recover its parameters. Finally, we empirically verify our theoretical findings in simple simulated domains. For instance, V * θ ∈ R S is the vector of optimal values, p θ (s, a) ∈ R S is the vector of transition probabilities from s, a, r θ ∈ R SA is the flattened reward matrix, and so on. To measure the distance between two models θ, θ , we define ∆ r s,a (θ, θ ) := |r θ (s, a)r θ (s, a)| for the rewards and ∆ p s,a (θ, θ ) = |(p θ (s, a)p θ (s, a)) T V * θ | for the transition probabilities. The latter measures how much the expected return of an agent taking s, a and acting optimally in θ changes when the first transition only is governed by θ . See Appendix A for a quick reference of notation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Enforcing Hard Constraints with Soft Barriers: Safe Reinforcement Learning in Unknown Stochastic EnvironmentsYixuan Wang, Simon Sinong Zhan, Ruochen Jiao, Zhilu Wang 等ICML 2023 · 被引用 81 次
- Near-Optimal Model-Free Reinforcement Learning in Non-Stationary Episodic MDPsWeichao Mao, Kaiqing Zhang, Ruihao Zhu, David Simchi-Levi 等ICML 2021 · 被引用 49 次
- Learning Mixtures of Linear Dynamical SystemsYanxi Chen, H. Vincent PoorICML 2022 · 被引用 22 次
- Learning in Non-Cooperative Configurable Markov Decision ProcessesGiorgia Ramponi, Alberto Maria Metelli, Alessandro Concetti, Marcello RestelliNeurIPS 2021 · 被引用 12 次
- CtrlFormer: Learning Transferable State Representation for Visual Control via TransformerYao Mark Mu, Shoufa Chen, Mingyu Ding, Jianyu Chen 等ICML 2022 · 被引用 10 次
它引用的顶会 Paper1
相关 Paper
- Policy Caches with Successor FeaturesMark W. Nemecek, Ron ParrICML 2021 · 被引用 19 次
- TempLe: Learning Template of Transitions for Sample Efficient Multi-task RLYanchao Sun, Xiangyu Yin, Furong HuangAAAI 2021 · 被引用 17 次
- Overcoming the Sim-to-Real Gap: Leveraging Simulation to Learn to Explore for Real-World RLAndrew Wagenmaker, Kevin Huang, Liyiming Ke, Kevin Jamieson 等NeurIPS 2024 · 被引用 45 次
- REPAINT: Knowledge Transfer in Deep Reinforcement LearningYunzhe Tao, Sahika Genc, Jonathan Chung, Tao Sun 等ICML 2021 · 被引用 32 次
- Provably sample-efficient RL with side information about latent dynamicsYao Liu, Dipendra Misra, Miro Dudík, Robert E. SchapireNeurIPS 2022 · 被引用 2 次
