Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement Learning
Aleksei Petrenko, Zhehui Huang, Tushar Kumar, Gaurav S. Sukhatme, Vladlen Koltun
摘要
Increasing the scale of reinforcement learning experiments has allowed researchers to achieve unprecedented results in both training sophisticated agents for video games, and in sim-to-real transfer for robotics. Typically such experiments rely on large distributed systems and require expensive hardware setups, limiting wider access to this exciting area of research. In this work we aim to solve this problem by optimizing the efficiency and resource utilization of reinforcement learning algorithms instead of relying on distributed computation. We present the "Sample Factory", a high-throughput training system optimized for a single-machine setting. Our architecture combines a highly efficient, asynchronous, GPU-based sampler with off-policy correction techniques, allowing us to achieve throughput higher than 10 5 environment frames/second on non-trivial control problems in 3D without sacrificing sample efficiency. We extend Sample Factory to support self-play and population-based training and apply these techniques to train highly capable agents for a multiplayer first-person shooter game. Github: https://github.com/ alex-petrenko/sample-factory
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans 等NeurIPS 2021 · 被引用 826 次
- Motif: Intrinsic Motivation from Artificial Intelligence FeedbackMartin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu 等ICLR 2024 · 被引用 97 次
- Accelerating Reinforcement Learning through GPU Atari EmulationSteven Dalton, Iuri FrosioNeurIPS 2020 · 被引用 50 次
- Large Batch Simulation for Deep Reinforcement LearningBrennan Shacklett, Erik Wijmans, Aleksei Petrenko, Manolis Savva 等ICLR 2021 · 被引用 29 次
- Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation ProblemMaciej Wolczyk, Bartlomiej Cupial, Mateusz Ostaszewski, Michal Bortkiewicz 等ICML 2024 · 被引用 29 次
它引用的顶会 Paper4
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu 等ICLR 2020 · 被引用 751 次
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee 等ICLR 2020 · 被引用 608 次
- Off-Policy Actor-Critic with Shared Experience ReplaySimon Schmitt, Matteo Hessel, Karen SimonyanICML 2020 · 被引用 71 次
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central InferenceLasse Espeholt, Raphaël Marinier, Piotr Stanczyk, Ke Wang 等ICLR 2020 · 被引用 32 次
相关 Paper
- Staggered Environment Resets Improve Massively Parallel On-Policy Reinforcement LearningSid Bharthulwar, Stone Tao, Hao SuNeurIPS 2025 · 被引用 6 次
- Parallel Q-Learning: Scaling Off-policy Reinforcement Learning under Massively Parallel SimulationZechu Li, Tao Chen, Zhang-Wei Hong, Anurag Ajay 等ICML 2023 · 被引用 27 次
- SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand CoresZhiyu Mei, Wei Fu, Jiaxuan Gao, Guangju Wang 等ICLR 2024 · 被引用 10 次
- Cheaper and Faster: Distributed Deep Reinforcement Learning with Serverless ComputingHanfei Yu, Jian Li, Yang Hua, Xu Yuan 等AAAI 2024 · 被引用 8 次
- VER: Scaling On-Policy RL Leads to the Emergence of Navigation in Embodied RearrangementErik Wijmans, Irfan Essa, Dhruv BatraNeurIPS 2022 · 被引用 24 次
