Domain Randomization via Entropy Maximization
Gabriele Tiboni, Pascal Klink, Jan Peters, Tatiana Tommasi, Carlo D'Eramo, Georgia Chalvatzaki
摘要
Varying dynamics parameters in simulation is a popular Domain Randomization (DR) approach for overcoming the reality gap in Reinforcement Learning (RL). Nevertheless, DR heavily hinges on the choice of the sampling distribution of the dynamics parameters, since high variability is crucial to regularize the agent's behavior but notoriously leads to overly conservative policies when randomizing excessively. In this paper, we propose a novel approach to address sim-to-real transfer, which automatically shapes dynamics distributions during training in simulation without requiring real-world data. We introduce DOmain RAndomization via Entropy MaximizatiON (DORAEMON), a constrained optimization problem that directly maximizes the entropy of the training distribution while retaining generalization capabilities. In achieving this, DORAEMON gradually increases the diversity of sampled dynamics parameters as long as the probability of success of the current policy is sufficiently high. We empirically validate the consistent benefits of DORAEMON in obtaining highly adaptive and generalizable policies, i.e. solving the task at hand across the widest range of dynamics parameters, as opposed to representative baselines from the DR literature. Notably, we also demonstrate the Sim2Real applicability of DORAEMON through its successful zero-shot transfer in a robotic manipulation setup under unknown real-world parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Rapidly Adapting Policies to the Real-World via Simulation-Guided Fine-TuningPatrick Yin, Tyler Westenbroek, Ching-An Cheng, Andrey Kolobov 等ICLR 2025
- On Generalization Across Environments In Multi-Objective Reinforcement LearningJayden Teoh, Pradeep Varakantham, Peter VamplewICLR 2025
- Flow-based Domain Randomization for Learning and Sequencing Robotic SkillsAidan Curtis, Eric Li, Michael Noseworthy, Nishad Gothoskar 等ICML 2025
- Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement LearningChi Zhang, Zi-Jia Wang, George K. Atia, Sihong He 等ICML 2025
- Fast and Robust: Task Sampling with Posterior and Diversity Synergies for Adaptive Decision-Makers in Randomized EnvironmentsYun Qu, Cheems Wang, Yixiu Mao, Yiqin Lv 等ICML 2025
它引用的顶会 Paper5
- Understanding Domain Randomization for Sim-to-real TransferXiaoyu Chen, Jiachen Hu, Chi Jin, Lihong Li 等ICLR 2022 · 被引用 164 次
- Self-Paced Deep Reinforcement LearningPascal Klink, Carlo D'Eramo, Jan Peters, Joni PajarinenNeurIPS 2020 · 被引用 83 次
- Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain AdaptationPeide Huang, Mengdi Xu, Jiacheng Zhu, Laixi Shi 等NeurIPS 2022 · 被引用 44 次
- Curriculum Reinforcement Learning via Constrained Optimal TransportPascal Klink, Haoyi Yang, Carlo D'Eramo, Jan Peters 等ICML 2022 · 被引用 44 次
- Outcome-directed Reinforcement Learning by Uncertainty & Temporal Distance-Aware Curriculum Goal GenerationDaesol Cho, Seungjae Lee, H. Jin KimICLR 2023
相关 Paper
- Statistical Guarantees for Offline Domain RandomizationArnaud Fickinger, Abderrahim Bendahi, Stuart RussellICLR 2026
- SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real TransferYarden As, Chengrui Qu, Benjamin Unger, Dongho Kang 等NeurIPS 2025 · 被引用 9 次
- Revisiting Domain Randomization via Relaxed State-Adversarial Policy OptimizationYun-Hsuan Lien, Ping-Chun Hsieh, Yu-Shuen WangICML 2023 · 被引用 1 次
- Monotonic Robust Policy Optimization with Model DiscrepancyYuankun Jiang, Chenglin Li, Wenrui Dai, Junni Zou 等ICML 2021 · 被引用 24 次
- Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real DeploymentGengyue Han, Yiheng FengICML 2026
