Harnessing Structures for Value-Based Planning and Reinforcement Learning
Yuzhe Yang, Guo Zhang, Zhi Xu, Dina Katabi
摘要
Value-based methods constitute a fundamental methodology in planning and deep reinforcement learning (RL). In this paper, we propose to exploit the underlying structures of the state-action value function, i.e., Q function, for both planning and deep RL. In particular, if the underlying system dynamics lead to some global structures of the Q function, one should be capable of inferring the function better by leveraging such structures. Specifically, we investigate the low-rank structure, which widely exists for big data matrices. We verify empirically the existence of low-rank Q functions in the context of control and deep RL tasks. As our key contribution, by leveraging Matrix Estimation (ME) techniques, we propose a general framework to exploit the underlying low-rank structure in Q functions. This leads to a more efficient planning procedure for classical control, and additionally, a simple scheme that can be applied to any value-based RL techniques to consistently achieve better performance on "low-rank" tasks. Extensive experiments on control tasks and Atari games confirm the efficacy of our approach. Code is available at this https URL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Implicit Under-Parameterization Inhibits Data-Efficient Deep Reinforcement LearningAviral Kumar, Rishabh Agarwal, Dibya Ghosh, Sergey LevineICLR 2021 · 被引用 155 次
- Sample Efficient Reinforcement Learning via Low-Rank Matrix EstimationDevavrat Shah, Dogyoon Song, Zhi Xu, Yuzhe YangNeurIPS 2020 · 被引用 35 次
- No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPOSkander Moalla, Andrea Miele, Daniil Pyatko, Razvan Pascanu 等NeurIPS 2024 · 被引用 35 次
- Simplicial Embeddings Improve Sample Efficiency in Actor–Critic AgentsJohan Obando-Ceron, Walter Mayor, Samuel Lavoie, Scott Fujimoto 等ICLR 2026 · 被引用 12 次
- Shift Before You Learn: Enabling Low-Rank Representations in Reinforcement LearningBastien Dubail, Stefan Stojanovic, Alexandre ProutièreNeurIPS 2025 · 被引用 4 次
相关 Paper
- Latent Variable Representation for Reinforcement LearningTongzheng Ren, Chenjun Xiao, Tianjun Zhang, Na Li 等ICLR 2023 · 被引用 1 次
- Spectral Entry-wise Matrix Estimation for Low-Rank Reinforcement LearningStefan Stojanovic, Yassir Jedra, Alexandre ProutièreNeurIPS 2023 · 被引用 9 次
- Foresee then Evaluate: Decomposing Value Estimation with Latent Future PredictionHongyao Tang, Zhaopeng Meng, Guangyong Chen, Pengfei Chen 等AAAI 2021 · 被引用 5 次
- Accelerating Q-learning through Efficient Value-sharing across ActionsPrabhat Nagarajan, Brett Daley, Martha White, Marlos C. MachadoICML 2026
- Actor-Free Continuous Control via Structurally Maximizable Q-FunctionsYigit Korkmaz, Urvi Bhuwania, Ayush Jain, Erdem BiyikNeurIPS 2025
