Scaling Offline RL via Efficient and Expressive Shortcut Models
Nicolas A. Espinosa Dice, Yiyi Zhang, Yiding Chen, Bradley Guo, Owen Oertell, Gokul Swamy, Kianté Brantley, Wen Sun
Abstract
Diffusion and flow models have emerged as powerful generative approaches capable of modeling diverse and multimodal behavior. However, applying these models to offline reinforcement learning (RL) remains challenging due to the iterative nature of their noise sampling processes, making policy optimization difficult. In this paper, we introduce Scalable Offline Reinforcement Learning (SORL), a new offline RL algorithm that leverages shortcut models - a novel class of generative models - to scale both training and inference. SORL's policy can capture complex data distributions and can be trained simply and efficiently in a one-stage training procedure. At test time, SORL introduces both sequential and parallel inference scaling by using the learned Q-function as a verifier. We demonstrate that SORL achieves strong performance across a range of offline RL tasks and exhibits positive scaling behavior with increased test-time compute. We release the code at nico-espinosadice.github.io/projects/sorl.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bbd6dbad-b9f4-4c6b-938b-5e62744dd290Cited by top-tier papers8
- Q-Learning with Adjoint MatchingQiyang Li, Sergey LevineICLR 2026 · 36 citations
- floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RLBhavya Agrawalla, Michal Nauman, Khush Agrawal, Aviral KumarICLR 2026 · 22 citations
- One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlowZeyuan Wang, Da Li, Yulin Chen, Ye Shi et al.AAAI 2026 · 6 citations
- ReFORM: Reflected Flows for On-support Offline RL via Noise ManipulationSongyuan Zhang, Oswin So, H. M. Sabbir Ahmad, Eric Yang Yu et al.ICLR 2026 · 5 citations
- Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-LearningThanh Nguyen, Tri Ton, Hongbin Choe, Minh-Tung Luu et al.ICML 2026 · 3 citations
Builds on51
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Flow-Based Single-Step Completion for Efficient and Expressive Policy LearningPrajwal Koirala, Cody FlemingICLR 2026 · 12 citations
- Score Regularized Policy Optimization through Diffusion BehaviorHuayu Chen, Cheng Lu, Zhengyi Wang, Hang Su et al.ICLR 2024 · 59 citations
- Diffusion Policies as an Expressive Policy Class for Offline Reinforcement LearningZhendong Wang, Jonathan J. Hunt, Mingyuan ZhouICLR 2023 · 33 citations
- Flow Q-LearningSeohong Park, Qiyang Li, Sergey LevineICML 2025
- Offline Reinforcement Learning via High-Fidelity Generative Behavior ModelingHuayu Chen, Cheng Lu, Chengyang Ying, Hang Su et al.ICLR 2023 · 6 citations
