SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement Learning
Shuai Zhang, Heshan Devaka Fernando, Miao Liu, Keerthiram Murugesan, Songtao Lu, Pin-Yu Chen, Tianyi Chen, Meng Wang
摘要
This paper studies the transfer reinforcement learning (RL) problem where multiple RL problems have different reward functions but share the same underlying transition dynamics. In this setting, the Q-function of each RL problem (task) can be decomposed into a successor feature (SF) and a reward mapping: the former characterizes the transition dynamics, and the latter characterizes the task-specific reward function. This Q-function decomposition, coupled with a policy improvement operator known as generalized policy improvement (GPI), reduces the sample complexity of finding the optimal Q-function, and thus the SF &GPI framework exhibits promising empirical performance compared to traditional RL methods like Q-learning. However, its theoretical foundations remain largely unestablished, especially when learning the successor features using deep neural networks (SF-DQN). This paper studies the provable knowledge transfer using SFs-DQN in transfer RL problems. We establish the first convergence analysis with provable generalization guarantees for SF-DQN with GPI. The theory reveals that SF-DQN with GPI outperforms conventional RL approaches, such as deep Q-network, in terms of both faster convergence rate and better generalization. Numerical experiments on real and synthetic RL tasks support the superior performance of SF-DQN &GPI, aligning with our theoretical findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Toward Understanding the Feature Learning Process of Self-supervised Contrastive LearningZixin Wen, Yuanzhi LiICML 2021 · 被引用 162 次
- Feature Purification: How Adversarial Training Performs Robust Deep LearningZeyuan Allen-Zhu, Yuanzhi LiFOCS 2021 · 被引用 83 次
- A Finite-Time Analysis of Q-Learning with Neural Network Function ApproximationPan Xu, Quanquan GuICML 2020 · 被引用 79 次
- A Theoretical Analysis on Feature Learning in Neural Networks: Emergence from Inputs and Advantage over Fixed FeaturesZhenmei Shi, Junyi Wei, Yingyu LiangICLR 2022 · 被引用 58 次
- Patch-level Routing in Mixture-of-Experts is Provably Sample-efficient for Convolutional Neural NetworksMohammed Nowaz Rabbani Chowdhury, Shuai Zhang, Meng Wang, Sijia Liu 等ICML 2023 · 被引用 45 次
相关 Paper
- Composing Task Knowledge With Modular Successor Feature ApproximatorsWilka Carvalho, Angelos Filos, Richard L. Lewis, Honglak Lee 等ICLR 2023 · 被引用 2 次
- A New Representation of Successor Features for Transfer across Dissimilar EnvironmentsMajid Abdolshah, Hung Le, Thommen George Karimpanal, Sunil Gupta 等ICML 2021 · 被引用 21 次
- Optimistic Linear Support and Successor Features as a Basis for Optimal Policy TransferLucas Nunes Alegre, Ana L. C. Bazzan, Bruno C. da SilvaICML 2022 · 被引用 36 次
- Risk-Aware Transfer in Reinforcement Learning using Successor FeaturesMichael Gimelfarb, André Barreto, Scott Sanner, Chi-Guhn LeeNeurIPS 2021 · 被引用 25 次
- Policy Caches with Successor FeaturesMark W. Nemecek, Ron ParrICML 2021 · 被引用 19 次
