RL-CycleGAN: Reinforcement Learning Aware Simulation-to-Real
Kanishka Rao, Chris Harris, Alex Irpan, Sergey Levine, Julian Ibarz, Mohi Khansari
摘要
Deep neural network based reinforcement learning (RL) can learn appropriate visual representations for complex tasks like vision-based robotic grasping without the need for manually engineering or prior learning a perception system. However, data for RL is collected via running an agent in the desired environment, and for applications like robotics, running a robot in the real world may be extremely costly and time consuming. Simulated training offers an appealing alternative, but ensuring that policies trained in simulation can transfer effectively into the real world requires additional machinery. Simulations may not match reality, and typically bridging the simulation-to-reality gap requires domain knowledge and task-specific engineering. We can automate this process by employing generative models to translate simulated images into realistic ones. However, this sort of translation is typically task-agnostic, in that the translated images may not preserve all features that are relevant to the task. In this paper, we introduce the RL-scene consistency loss for image translation, which ensures that the translation operation is invariant with respect to the Q-values associated with the image. This allows us to learn a task-aware translation. Incorporating this loss into unsupervised domain translation, we obtain the RL-CycleGAN, a new approach for simulation-to-real-world transfer for reinforcement learning. In evaluations of RL-CycleGAN on two vision-based robotics grasping tasks, we show that RL-CycleGAN offers a substantial improvement over a number of prior methods for sim-to-real transfer, attaining excellent real-world performance with only a modest number of real-world observations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated ObjectsRuihai Wu, Yan Zhao, Kaichun Mo, Zizheng Guo 等ICLR 2022 · 被引用 119 次
- Should I Run Offline Reinforcement Learning or Behavioral Cloning?Aviral Kumar, Joey Hong, Anikait Singh, Sergey LevineICLR 2022 · 被引用 84 次
- Sketch Your Own GANSheng-Yu Wang, David Bau, Jun-Yan ZhuICCV 2021 · 被引用 82 次
- When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement LearningHaoyi Niu, Shubham Sharma, Yiwen Qiu, Ming Li 等NeurIPS 2022 · 被引用 81 次
- Learning Cross-Domain Correspondence for Control with Dynamics Cycle-ConsistencyQiang Zhang, Tete Xiao, Alexei A. Efros, Lerrel Pinto 等ICLR 2021 · 被引用 73 次
相关 Paper
- Domain Adaptation In Reinforcement Learning Via Latent Unified State RepresentationJinwei Xing, Takashi Nagata, Kexin Chen, Xinyun Zou 等AAAI 2021 · 被引用 65 次
- Deep CG2Real: Synthetic-to-Real Translation via Image DisentanglementSai Bi, Kalyan Sunkavalli, Federico Perazzi, Eli Shechtman 等ICCV 2019 · 被引用 37 次
- StereoGAN: Bridging Synthetic-to-Real Domain Gap by Joint Optimization of Domain Translation and Stereo MatchingRui Liu, Chengxi Yang, Wenxiu Sun, Xiaogang Wang 等CVPR 2020
- Kernel of CycleGAN as a principal homogeneous spaceNikita Moriakov, Jonas Adler, Jonas TeuwenICLR 2020 · 被引用 6 次
- On Translation and Reconstruction Guarantees of the Cycle-Consistent Generative Adversarial NetworksAnish Chakrabarty, Swagatam DasNeurIPS 2022 · 被引用 5 次
