SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Modeling
Yixian Zhang, Shu'ang Yu, Tonghe Zhang, Mo Guang, Haojia Hui, Kaiwen Long, Yu Wang, Chao Yu, Wenbo Ding
摘要
Training expressive flow-based policies with off-policy reinforcement learning is notoriously unstable due to gradient pathologies in the multi-step action sampling process. We trace this instability to a fundamental connection: the flow rollout is algebraically equivalent to a residual recurrent computation, making it susceptible to the same vanishing and exploding gradients as RNNs. To address this, we reparameterize the velocity network using principles from modern sequential models, introducing two stable architectures: Flow-G, which incorporates a gated velocity, and Flow-T, which utilizes a decoded velocity. We then develop a practical SAC-based algorithm, enabled by a noise-augmented rollout, that facilitates direct end-to-end training of these policies. Our approach supports both from-scratch and offline-to-online learning and achieves state-of-the-art performance on continuous control and robotic manipulation benchmarks, eliminating the need for common workarounds like policy distillation or surrogate objectives. Anonymized code is available at https://anonymous.4open.science/r/SAC-FLOW
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- RLinf: Flexible and Efficient Large-Scale Reinforcement Learning via Macro-to-Micro Flow TransformationChao Yu, Yuanqing Wang, Zhen Guo, Hao Lin 等OSDI 2026 · 被引用 63 次
- FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge MatchingLei Lyu, Yunfei Li, Yu Luo, Fuchun Sun 等ICML 2026 · 被引用 5 次
- Causal Flow Q-Learning for Robust Offline Reinforcement LearningMingxuan Li, Junzhe Zhang, Elias BareinboimICML 2026 · 被引用 1 次
- Generative Online Reinforcement LearningChubin Zhang, Zhenglin Wan, Feng Chen, Fuchao Yang 等ICML 2026 · 被引用 1 次
- Reparameterization Flow Policy OptimizationHai Zhong, Zhuoran Li, Xun Wang, Longbo HuangICML 2026
它引用的顶会 Paper17
- WCSAC: Worst-Case Soft Actor Critic for Safety-Constrained Reinforcement LearningQisong Yang, Thiago D. Simão, Simon H. Tindemans, Matthijs T. J. SpaanAAAI 2021 · 被引用 168 次
- Data Quality in Imitation LearningSuneel Belkhale, Yuchen Cui, Dorsa SadighNeurIPS 2023 · 被引用 135 次
- Diffusion-based Reinforcement Learning via Q-weighted Variational Policy OptimizationShutong Ding, Ke Hu, Zhenhao Zhang, Kan Ren 等NeurIPS 2024 · 被引用 132 次
- Reinforcement Learning with Action ChunkingQiyang Li, Zhiyuan Zhou, Sergey LevineNeurIPS 2025 · 被引用 114 次
- ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement LearningTonghe Zhang, Chao Yu, Sichang Su, Yu WangNeurIPS 2025 · 被引用 101 次
相关 Paper
- One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlowZeyuan Wang, Da Li, Yulin Chen, Ye Shi 等AAAI 2026 · 被引用 6 次
- Bridging Successor Measure and Online Policy Learning with Flow Matching-Based RepresentationsHaosen Shi, Jianda Chen, Sinno Jialin PanICLR 2026
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform SamplingChe Wang, Yanqiu Wu, Quan Vuong, Keith W. RossICML 2020 · 被引用 38 次
- PolicyFlow: Policy Optimization with Continuous Normalizing Flow in Reinforcement LearningShunpeng Yang, Ben Liu, Hua ChenICLR 2026 · 被引用 6 次
- Q-Flow: Stable and Expressive Reinforcement Learning with Flow-based PolicyJaeHyeok Doo, Byeongguk Jeon, Seonghyeon Ye, Kimin Lee 等ICML 2026 · 被引用 1 次
