Rapidly Adapting Policies to the Real-World via Simulation-Guided Fine-Tuning
Patrick Yin, Tyler Westenbroek, Ching-An Cheng, Andrey Kolobov, Abhishek Gupta
Abstract
Robot learning requires a considerable amount of high-quality data to realize the promise of generalization. However, large data sets are costly to collect in the real world. Physics simulators can cheaply generate vast data sets with broad coverage over states, actions, and environments. However, physics engines are fundamentally misspecified approximations to reality. This makes direct zero-shot transfer from simulation to reality challenging, especially in tasks where precise and force-sensitive manipulation is necessary. Thus, fine-tuning these policies with small real-world data sets is an appealing pathway for scaling robot learning. However, current reinforcement learning fine-tuning frameworks leverage general, unstructured exploration strategies which are too inefficient to make realworld adaptation practical. This paper introduces the Simulation-Guided Finetuning (SGFT) framework, which demonstrates how to extract structural priors from physics simulators to substantially accelerate real-world adaptation. Specifically, our approach uses a value function learned in simulation to guide real-world exploration. We demonstrate this approach across five real-world dexterous manipulation tasks where zero-shot sim-to-real transfer fails. We further demonstrate our framework substantially outperforms baseline fine-tuning methods, requiring up to an order of magnitude fewer real-world samples and succeeding at difficult tasks where prior approaches fail entirely. Last but not least, we provide theoretical justification for this new paradigm which underpins how SGFT can rapidly learn high-performance policies in the face of large sim-to-real dynamics gaps. Project webpage: weirdlabuw.github.io/sgft
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0efa61c6-930a-4c91-b8fa-10658d121865Cited by top-tier papers1
Ask how each one uses itBuilds on22
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- 🏘️ ProcTHOR: Large-Scale Embodied AI Using Procedural GenerationMatt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs et al.NeurIPS 2022 · 596 citations
- Eureka: Human-Level Reward Design via Coding Large Language ModelsYecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang et al.ICLR 2024 · 582 citations
Related papers
- Contact-Aware Neural DynamicsChangwei Jing, Jai Krishna Bandi, Jianglong Ye, Yan Duan et al.CVPR 2026 · 5 citations
- ASID: Active Exploration for System Identification in Robotic ManipulationMarius Memmel, Andrew Wagenmaker, Chuning Zhu, Dieter Fox et al.ICLR 2024 · 36 citations
- Emergent Dexterity Via Diverse Resets and Large-Scale Reinforcement LearningPatrick Yin, Tyler Westenbroek, Zhengyu Zhang, Ignacio Dagnino et al.ICLR 2026 · 15 citations
- Can Vision Language Models Learn Intuitive Physics from Interaction?Luca M. Schulze Buschoff, Konstantinos Voudouris, Can Demircan, Eric SchulzICML 2026
- Latent Adaptation of Foundation Policies for Sim-to-Real TransferLongchao Da, Thirulogasankar Pranav Kutralingam, Lirong Xiang, Hua WeiICLR 2026
