Towards Bridging the Gap between Large-Scale Pretraining and Efficient Finetuning for Humanoid Control
Weidong Huang, Zhehan Li, Hangxin Liu, Biao Hou, Yao Su, Jingwen Zhang
摘要
Reinforcement Learning (RL) is widely used for humanoid control, with on-policy methods such as Proximal Policy Optimization (PPO) enabling robust training via large-scale parallel simulation and, in some cases, zero-shot deployment to real robots. However, the low sample efficiency of on-policy algorithms limits safe adaptation to new environments. Although off-policy RL and model-based RL have shown improved sample efficiency, the gap between large-scale pretraining and efficient finetuning on humanoids still exists. In this paper, we find that off-policy Soft Actor-Critic (SAC), with large-batch update and a high Update-To-Data (UTD) ratio, reliably supports large-scale pretraining of humanoid locomotion policies, achieving zero-shot deployment on real robots. For adaptation, we demonstrate that these SAC-pretrained policies can be finetuned in new environments and out-of-distribution tasks using model-based methods. Data collection in the new environment executes a deterministic policy while stochastic exploration is instead confined to a physics-informed world model. This separation mitigates the risks of random exploration during adaptation while preserving exploratory coverage for improvement. Overall, the approach couples the wall-clock efficiency of large-scale simulation during pretraining with the sample efficiency of model-based learning during fine-tuning. Code and videos: https://lift-humanoid.github.io * Corresponding Author.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- Uncertainty-Based Offline Reinforcement Learning with Diversified Q-EnsembleGaon An, Seungyong Moon, Jang-Hyun Kim, Hyun Oh SongNeurIPS 2021 · 被引用 430 次
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 被引用 326 次
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon 等ICML 2022 · 被引用 269 次
- Dropout Q-Functions for Doubly Efficient Reinforcement LearningTakuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi 等ICLR 2022 · 被引用 157 次
相关 Paper
- Can Agents Run Relay Race with Strangers? Generalization of RL to Out-of-Distribution TrajectoriesLi-Cheng Lan, Huan Zhang, Cho-Jui HsiehICLR 2023 · 被引用 3 次
- Latent Adaptation of Foundation Policies for Sim-to-Real TransferLongchao Da, Thirulogasankar Pranav Kutralingam, Lirong Xiang, Hua WeiICLR 2026
- Accelerated Policy Learning with Parallel Differentiable SimulationJie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos 等ICLR 2022 · 被引用 141 次
- FOSP: Fine-tuning Offline Safe Policy through World ModelsChenyang Cao, Yucheng Xin, Silang Wu, Longxiang He 等ICLR 2025
- Adaptive Horizon Actor-Critic for Policy Learning in Contact-Rich Differentiable SimulationIgnat Georgiev, Krishnan Srinivasan, Jie Xu, Eric Heiden 等ICML 2024 · 被引用 27 次
