Emergent Dexterity Via Diverse Resets and Large-Scale Reinforcement Learning
Patrick Yin, Tyler Westenbroek, Zhengyu Zhang, Ignacio Dagnino, Eeshani Shilamkar, Numfor Mbiziwo-Tiapo, Simran Bagaria, Xinlei Liu, Galen Mullins, Andrey Kolobov, Abhishek Gupta
摘要
Reinforcement learning in massively parallel physics simulations has driven major progress in sim-to-real robot learning. However, current approaches remain brittle and task-specific, relying on extensive per-task engineering to design rewards, curricula, and demonstrations. Even with this engineering, typical reinforcement learning methods can often fail on long-horizon, contact-rich manipulation tasks and do not meaningfully scale with compute, as performance quickly saturates when training revisits the same narrow regions of state space. We introduce OmniReset, a simple and scalable framework that enables on-policy reinforcement learning to robustly solve a broad class of dexterous manipulation tasks using fixed algorithm hyperparameters, no curricula, minimal reward engineering, and no human demonstrations. Our key insight is that long-horizon exploration can be dramatically simplified by using simulator resets to systematically expose the RL algorithm to the diverse set of robot-object interactions that underlie dexterous manipulation. OmniReset programmatically generates such resets with minimal human input, converting additional compute directly into broader behavioral coverage and continued performance gains for dynamic policies. We show that OmniReset gracefully scales to long-horizon dexterous manipulation tasks beyond the capabilities of existing approaches and is able to learn robust policies demonstrating a variety of dynamic, contact-rich recovery behavior. Finally, we distill OmniReset into visuomotor policies that can be transferred to the real world zero-shot, displaying robust retrying behavior to accomplish complex, contact-rich tasks with non-trivial success rates. Project webpage: https://omnireset.github.io
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 被引用 299 次
- Curious Exploration via Structured World Models Yields Zero-Shot Object ManipulationCansu Sancaktar, Sebastian Blaes, Georg MartiusNeurIPS 2022 · 被引用 43 次
- CCIL: Continuity-Based Data Augmentation for Corrective Imitation LearningLiyiming Ke, Yunchu Zhang, Abhay Deshpande, Siddhartha S. Srinivasa 等ICLR 2024 · 被引用 33 次
- Mastering the Unsupervised Reinforcement Learning Benchmark from PixelsSai Rajeswar, Pietro Mazzaglia, Tim Verbelen, Alexandre Piché 等ICML 2023 · 被引用 30 次
相关 Paper
- The Ingredients of Real World Robotic Reinforcement LearningHenry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah 等ICLR 2020 · 被引用 202 次
- Rapidly Adapting Policies to the Real-World via Simulation-Guided Fine-TuningPatrick Yin, Tyler Westenbroek, Ching-An Cheng, Andrey Kolobov 等ICLR 2025
- Visual Sim-to-Real at Scale for Humanoid Loco-ManipulationTairan He, Zi Wang, Haoru Xue, Qingwei Ben 等CVPR 2026
- Benchmarking Offline Reinforcement Learning on Real-Robot HardwareNico Gürtler, Sebastian Blaes, Pavel Kolev, Felix Widmaier 等ICLR 2023 · 被引用 11 次
- Staggered Environment Resets Improve Massively Parallel On-Policy Reinforcement LearningSid Bharthulwar, Stone Tao, Hao SuNeurIPS 2025 · 被引用 6 次
