GridToPix: Training Embodied Agents with Minimal Supervision
Unnat Jain, Iou-Jen Liu, Svetlana Lazebnik, Aniruddha Kembhavi, Luca Weihs, Alexander G. Schwing
摘要
While deep reinforcement learning (RL) promises freedom from hand-labeled data, great successes, especially for Embodied AI, require significant work to create supervision via carefully shaped rewards. Indeed, without shaped rewards, i.e., with only terminal rewards, present-day Embodied AI results degrade significantly across Embodied AI problems from single-agent Habitat-based PointGoal Navigation (SPL drops from 55 to 0) and two-agent AI2-THOR-based Furniture Moving (success drops from 58% to 1%) to three-agent Google Football-based 3 vs. 1 with Keeper (game score drops from 0.6 to 0.1). As training from shaped rewards doesn’t scale to more realistic tasks, the community needs to improve the success of training with terminal rewards. For this we propose GRIDTOPIX: 1) train agents with terminal rewards in gridworlds that generically mirror Embodied AI environments, i.e., they are independent of the task; 2) distill the learned policy into agents that reside in complex visual worlds. Despite learning from only terminal rewards with identical models and RL algorithms, GRIDTOPIX significantly improves results across tasks: from PointGoal Navigation (SPL improves from 0 to 64) and Furniture Moving (success improves from 1% to 25%) to football gameplay (game score improves from 0.1 to 0.6). GRIDTOPIX even helps to improve the results of shaped reward training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Habitat 3.0: A Co-Habitat for Humans, Avatars, and RobotsXavier Puig, Eric Undersander, Andrew Szot, Mikael Dallaire Cote 等ICLR 2024 · 被引用 252 次
- Simple but Effective: CLIP Embeddings for Embodied AIApoorv Khandelwal, Luca Weihs, Roozbeh Mottaghi, Aniruddha KembhaviCVPR 2022 · 被引用 149 次
- Bridging the Imitation Gap by Adaptive InsubordinationLuca Weihs, Unnat Jain, Iou-Jen Liu, Jordi Salvador 等NeurIPS 2021 · 被引用 53 次
- Pretrained Language Models as Visual Planners for Human AssistanceDhruvesh Patel, Hamid Eghbalzadeh, Nitin Kamra, Michael Louis Iuzzolino 等ICCV 2023 · 被引用 41 次
- Learning Active Camera for Multi-Object NavigationPeihao Chen, Dongyu Ji, Kunyang Lin, Weiwen Hu 等NeurIPS 2022 · 被引用 40 次
它引用的顶会 Paper22
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans 等NeurIPS 2021 · 被引用 826 次
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu 等ICLR 2020 · 被引用 751 次
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee 等ICLR 2020 · 被引用 608 次
- Google Research Football: A Novel Reinforcement Learning EnvironmentKarol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajac 等AAAI 2020 · 被引用 496 次
相关 Paper
- Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied NavigationJuncheng Li, Xin Wang, Siliang Tang, Haizhou Shi 等CVPR 2020
- ELIGN: Expectation Alignment as a Multi-Agent Intrinsic RewardZixian Ma, Rose E. Wang, Fei-Fei Li, Michael S. Bernstein 等NeurIPS 2022 · 被引用 22 次
- Cross-View Policy Learning for Street NavigationAng Li, Huiyi Hu, Piotr Mirowski, Mehrdad FarajtabarICCV 2019 · 被引用 35 次
- A Simple Approach for Visual Room Rearrangement: 3D Mapping and Semantic SearchBrandon Trabucco, Gunnar A. Sigurdsson, Robinson Piramuthu, Gaurav S. Sukhatme 等ICLR 2023
- Progressor: A Perceptually Guided Reward Estimator with Self-Supervised Online RefinementTewodros W. Ayalew, Xiao Zhang, Kevin Yuanbo Wu, Tianchong Jiang 等ICCV 2025 · 被引用 13 次
