RFS: Reinforcement learning with Residual flow steering for dexterous manipulation
Entong Su, Tyler Westenbroek, Anusha Nagabandi, Abhishek Gupta
摘要
Imitation learning has emerged as an effective approach for bootstrapping sequential decision-making in robotics, achieving strong performance even in high-dimensional dexterous manipulation tasks. Recent behavior cloning methods further leverage expressive generative models, such as diffusion models and flow matching, to represent multimodal action distributions. However, policies pretrained in this manner often exhibit limited generalization and require additional fine-tuning to achieve robust performance at deployment time. Such adaptation must preserve the global exploration benefits of pretraining while enabling rapid correction of local execution errors. We propose Residual Flow Steering (RFS), a data-efficient reinforcement learning framework for adapting pretrained generative policies. RFS steers a pretrained flow-matching policy by jointly optimizing a residual action and a latent noise distribution, enabling complementary forms of exploration: local refinement through residual corrections and global exploration through latent-space modulation. This design allows efficient adaptation while retaining the expressive structure of the pretrained policy. We demonstrate the effectiveness of RFS on dexterous manipulation tasks, showing efficient fine-tuning both in simulation and in real-world settings when adapting pretrained base policies. Project website: https://weirdlabuw.github.io/rfs/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 被引用 326 次
相关 Paper
- Lagrangian Perturbation Diffusion Steering: Latent Reinforcement Learning for Generative PoliciesHikmet Simsir, Ozgur S. OguzICML 2026 · 被引用 1 次
- Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent GuidanceYang Zhang, Chenwei Wang, Ouyang Lu, Yuan Zhao 等ICLR 2026 · 被引用 21 次
- From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL FinetuningZhanyi Sun, shuran songICML 2026 · 被引用 6 次
- FlowPolicy: Enabling Fast and Robust 3D Flow-Based Policy via Consistency Flow Matching for Robot ManipulationQinglun Zhang, Zhen Liu, Haoqiang Fan, Guanghui Liu 等AAAI 2025 · 被引用 5 次
- Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow ModelsHongyin Zhang, Shiyuan Zhang, Junxi Jin, Qixin Zeng 等AAAI 2026 · 被引用 11 次
