Sample Efficient Reinforcement Learning via Low-Rank Matrix Estimation
Devavrat Shah, Dogyoon Song, Zhi Xu, Yuzhe Yang
摘要
We consider the question of learning -function in a sample efficient manner for reinforcement learning with continuous state and action spaces under a generative model. If -function is Lipschitz continuous, then the minimal sample complexity for estimating -optimal -function is known to scale as per classical non-parametric learning theory, where and denote the dimensions of the state and action spaces respectively. The -function, when viewed as a kernel, induces a Hilbert-Schmidt operator and hence possesses square-summable spectrum. This motivates us to consider a parametric class of -functions parameterized by its "rank" , which contains all Lipschitz -functions as . As our key contribution, we develop a simple, iterative learning algorithm that finds -optimal -function with sample complexity of when the optimal -function has low rank and the discounting factor is below a certain threshold. Thus, this provides an exponential improvement in sample complexity. To enable our result, we develop a novel Matrix Estimation algorithm that faithfully estimates an unknown low-rank matrix in the sense even in the presence of arbitrary bounded noise, which might be of interest in its own right. Empirical results on several stochastic control tasks confirm the efficacy of our "low-rank" algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- From Self-Attention to Markov Models: Unveiling the Dynamics of Generative TransformersMuhammed Emrullah Ildiz, Yixiao Huang, Yingcong Li, Ankit Singh Rawat 等ICML 2024 · 被引用 45 次
- PerSim: Data-Efficient Offline Reinforcement Learning with Heterogeneous Agents via Personalized SimulatorsAnish Agarwal, Abdullah Omar Alomar, Varkey Alumootil, Devavrat Shah 等NeurIPS 2021 · 被引用 22 次
- Agnostic Reinforcement Learning with Low-Rank MDPs and Rich ObservationsAyush Sekhari, Christoph Dann, Mehryar Mohri, Yishay Mansour 等NeurIPS 2021 · 被引用 15 次
- Spectral Entry-wise Matrix Estimation for Low-Rank Reinforcement LearningStefan Stojanovic, Yassir Jedra, Alexandre ProutièreNeurIPS 2023 · 被引用 9 次
- Combining Explicit and Implicit Regularization for Efficient Learning in Deep NetworksDan ZhaoNeurIPS 2022 · 被引用 9 次
它引用的顶会 Paper2
相关 Paper
- Sample-Efficient Reinforcement Learning for Linearly-Parameterized MDPs with a Generative ModelBingyan Wang, Yuling Yan, Jianqing FanNeurIPS 2021 · 被引用 26 次
- Represent to Control Partially Observed Systems: Representation Learning with Provable Sample EfficiencyLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2023
- Model-free Low-Rank Reinforcement Learning via Leveraged Entry-wise Matrix EstimationStefan Stojanovic, Yassir Jedra, Alexandre ProutièreNeurIPS 2024 · 被引用 2 次
- Sample Efficient Reinforcement Learning with Partial Dynamics KnowledgeMeshal Alharbi, Mardavij Roozbehani, Munther A. DahlehAAAI 2024 · 被引用 4 次
- On the Global Convergence of Fitted Q-Iteration with Two-layer Neural Network ParametrizationMudit Gaur, Vaneet Aggarwal, Mridul AgarwalICML 2023 · 被引用 3 次
