On the Power of Pre-training for Generalization in RL: Provable Benefits and Hardness
Haotian Ye, Xiaoyu Chen, Liwei Wang, Simon Shaolei Du
摘要
Generalization in Reinforcement Learning (RL) aims to train an agent during training that generalizes to the target environment. In this work, we first point out that RL generalization is fundamentally different from the generalization in supervised learning, and fine-tuning on the target environment is necessary for good test performance. Therefore, we seek to answer the following question: how much can we expect pretraining over training environments to be helpful for efficient and effective fine-tuning? On one hand, we give a surprising result showing that asymptotically, the improvement from pretraining is at most a constant factor. On the other hand, we show that pre-training can be indeed helpful in the non-asymptotic regime by designing a policy collection-elimination (PCE) algorithm and proving a distribution-dependent regret bound that is independent of the state-action space. We hope our theoretical results can provide insight towards understanding pre-training and generalization in RL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Overcoming the Sim-to-Real Gap: Leveraging Simulation to Learn to Explore for Real-World RLAndrew Wagenmaker, Kevin Huang, Liyiming Ke, Kevin Jamieson 等NeurIPS 2024 · 被引用 45 次
- Test-Time Regret Minimization in Meta Reinforcement LearningMirco Mutti, Aviv TamarICML 2024 · 被引用 4 次
- Constrained Meta Reinforcement Learning with Provable Test-Time SafetyTingting Ni, Maryam KamgarpourICML 2026
- On the Generalization Ability of Next-Token-Prediction PretrainingZhihao Li, Xue Jiang, Liyuan Liu, Xuelin Zhang 等ICML 2025
- Provable Zero-Shot Generalization in Offline Reinforcement LearningZhiyong Wang, Chen Yang, John C. S. Lui, Dongruo ZhouICML 2025
它引用的顶会 Paper15
- FLAMBE: Structural Complexity and Representation Learning of Low Rank MDPsAlekh Agarwal, Sham M. Kakade, Akshay Krishnamurthy, Wen SunNeurIPS 2020 · 被引用 271 次
- Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial ObservabilityDibya Ghosh, Jad Rahme, Aviral Kumar, Amy Zhang 等NeurIPS 2021 · 被引用 176 次
- Reinforcement Learning with General Value Function Approximation: Provably Efficient Approach via Bounded Eluder DimensionRuosong Wang, Ruslan Salakhutdinov, Lin F. YangNeurIPS 2020 · 被引用 168 次
- RL for Latent MDPs: Regret Guarantees and a Lower BoundJeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, Shie MannorNeurIPS 2021 · 被引用 91 次
- Sample-Efficient Reinforcement Learning of Undercomplete POMDPsChi Jin, Sham M. Kakade, Akshay Krishnamurthy, Qinghua LiuNeurIPS 2020 · 被引用 88 次
相关 Paper
- Towards Provable Emergence of In-Context Reinforcement LearningJiuqi Wang, Rohan Chandra, Shangtong ZhangNeurIPS 2025 · 被引用 5 次
- How Ensembles of Distilled Policies Improve Generalisation in Reinforcement LearningMax Weltevrede, Moritz A. Zanger, Matthijs T. J. Spaan, Wendelin BoehmerNeurIPS 2025 · 被引用 1 次
- Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline DataZhiyuan Zhou, Andy Peng, Qiyang Li, Sergey Levine 等ICLR 2025
- Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RLQin-Wen Luo, Ming-Kun Xie, Ye-Wen Wang, Sheng-Jun HuangNeurIPS 2024 · 被引用 15 次
- Investigating Pre-Training Objectives for Generalization in Vision-Based Reinforcement LearningDonghu Kim, Hojoon Lee, Kyungmin Lee, Dongyoon Hwang 等ICML 2024 · 被引用 3 次
