RL-CycleGAN: Reinforcement Learning Aware Simulation-to-Real
Kanishka Rao, Chris Harris, Alex Irpan, Sergey Levine, Julian Ibarz, Mohi Khansari
Abstract
Deep neural network based reinforcement learning (RL) can learn appropriate visual representations for complex tasks like vision-based robotic grasping without the need for manually engineering or prior learning a perception system. However, data for RL is collected via running an agent in the desired environment, and for applications like robotics, running a robot in the real world may be extremely costly and time consuming. Simulated training offers an appealing alternative, but ensuring that policies trained in simulation can transfer effectively into the real world requires additional machinery. Simulations may not match reality, and typically bridging the simulation-to-reality gap requires domain knowledge and task-specific engineering. We can automate this process by employing generative models to translate simulated images into realistic ones. However, this sort of translation is typically task-agnostic, in that the translated images may not preserve all features that are relevant to the task. In this paper, we introduce the RL-scene consistency loss for image translation, which ensures that the translation operation is invariant with respect to the Q-values associated with the image. This allows us to learn a task-aware translation. Incorporating this loss into unsupervised domain translation, we obtain the RL-CycleGAN, a new approach for simulation-to-real-world transfer for reinforcement learning. In evaluations of RL-CycleGAN on two vision-based robotics grasping tasks, we show that RL-CycleGAN offers a substantial improvement over a number of prior methods for sim-to-real transfer, attaining excellent real-world performance with only a modest number of real-world observations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated ObjectsRuihai Wu, Yan Zhao, Kaichun Mo, Zizheng Guo et al.ICLR 2022 · 119 citations
- Should I Run Offline Reinforcement Learning or Behavioral Cloning?Aviral Kumar, Joey Hong, Anikait Singh, Sergey LevineICLR 2022 · 84 citations
- Sketch Your Own GANSheng-Yu Wang, David Bau, Jun-Yan ZhuICCV 2021 · 82 citations
- When to Trust Your Simulator: Dynamics-Aware Hybrid Offline-and-Online Reinforcement LearningHaoyi Niu, Shubham Sharma, Yiwen Qiu, Ming Li et al.NeurIPS 2022 · 81 citations
- Learning Cross-Domain Correspondence for Control with Dynamics Cycle-ConsistencyQiang Zhang, Tete Xiao, Alexei A. Efros, Lerrel Pinto et al.ICLR 2021 · 73 citations
Related papers
- Domain Adaptation In Reinforcement Learning Via Latent Unified State RepresentationJinwei Xing, Takashi Nagata, Kexin Chen, Xinyun Zou et al.AAAI 2021 · 65 citations
- Deep CG2Real: Synthetic-to-Real Translation via Image DisentanglementSai Bi, Kalyan Sunkavalli, Federico Perazzi, Eli Shechtman et al.ICCV 2019 · 37 citations
- StereoGAN: Bridging Synthetic-to-Real Domain Gap by Joint Optimization of Domain Translation and Stereo MatchingRui Liu, Chengxi Yang, Wenxiu Sun, Xiaogang Wang et al.CVPR 2020
- Kernel of CycleGAN as a principal homogeneous spaceNikita Moriakov, Jonas Adler, Jonas TeuwenICLR 2020 · 6 citations
- On Translation and Reconstruction Guarantees of the Cycle-Consistent Generative Adversarial NetworksAnish Chakrabarty, Swagatam DasNeurIPS 2022 · 5 citations
