Towards General-Purpose Model-Free Reinforcement Learning
Scott Fujimoto, Pierluca D'Oro, Amy Zhang, Yuandong Tian, Michael Rabbat
摘要
Reinforcement learning (RL) promises a framework for near-universal problem-solving. In practice however, RL algorithms are often tailored to specific benchmarks, relying on carefully tuned hyperparameters and algorithmic choices. Recently, powerful model-based RL methods have shown impressive general results across benchmarks but come at the cost of increased complexity and slow run times, limiting their broader applicability. In this paper, we attempt to find a unifying model-free deep RL algorithm that can address a diverse class of domains and problem settings. To achieve this, we leverage model-based representations that approximately linearize the value function, taking advantage of the denser task objectives used by model-based RL while avoiding the costs associated with planning or simulated trajectories. We evaluate our algorithm, MR.Q, on a variety of common RL benchmarks with a single set of hyperparameters and show a competitive performance against domain-specific and general baselines, providing a concrete step towards building general-purpose model-free deep RL algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- XQC: Well-conditioned Optimization Accelerates Deep Reinforcement LearningDaniel Palenicek, Florian Vogt, Joe Watson, Ingmar Posner 等ICLR 2026 · 被引用 20 次
- TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement LearningMarco Bagatella, Matteo Pirotta, Ahmed Touati, Alessandro Lazaric 等ICLR 2026 · 被引用 19 次
- Simplicial Embeddings Improve Sample Efficiency in Actor–Critic AgentsJohan Obando-Ceron, Walter Mayor, Samuel Lavoie, Scott Fujimoto 等ICLR 2026 · 被引用 12 次
- Dynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot GeneralizationFrank Röder, Jan Benad, Manfred Eppe, Pradeep Kr. BanerjeeNeurIPS 2025 · 被引用 9 次
- Relative Entropy Pathwise Policy OptimizationClaas Voelcker, Axel Brunnbauer, Marcel Hussing, Michal Nauman 等ICLR 2026 · 被引用 6 次
它引用的顶会 Paper37
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann 等ICML 2020 · 被引用 584 次
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 被引用 457 次
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 被引用 437 次
- Data-Efficient Reinforcement Learning with Self-Predictive RepresentationsMax Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm 等ICLR 2021 · 被引用 399 次
相关 Paper
- QORA: Zero-Shot Transfer via Interpretable Object-Relational Model LearningGabriel Stella, Dmitri LoguinovICML 2024 · 被引用 1 次
- MAMBA: an Effective World Model Approach for Meta-Reinforcement LearningZohar Rimon, Tom Jurgenson, Orr Krupnik, Gilad Adler 等ICLR 2024 · 被引用 15 次
- MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-ScenariosXuantang Xiong, Ni Mu, Runpeng Xie, Senhao Yang 等AAAI 2026
- A Clean Slate for Offline Reinforcement LearningMatthew Thomas Jackson, Uljad Berdica, Jarek Liesen, Shimon Whiteson 等NeurIPS 2025 · 被引用 11 次
- Towards an Information Theoretic Framework of Context-Based Offline Meta-Reinforcement LearningLanqing Li, Hai Zhang, Xinyu Zhang, Shatong Zhu 等NeurIPS 2024 · 被引用 24 次
