Compliant Residual DAgger: Improving Real-World Contact-Rich Manipulation with Human Corrections
Xiaomeng Xu, Yifan Hou, Zeyi Liu, Shuran Song
摘要
We address key challenges in Dataset Aggregation (DAgger) for real-world contactrich manipulation: how to collect informative human correction data and how to effectively update policies with this new data. We introduce Compliant Residual DAgger (CR-DAgger), which contains two novel components: 1) a Compliant Intervention Interface that leverages compliance control, allowing humans to provide gentle, accurate delta action corrections without interrupting the ongoing robot policy execution; and 2) a Compliant Residual Policy formulation that learns from human corrections while incorporating force feedback and force control.
Our system significantly enhances performance on precise contact-rich manipulation tasks using minimal correction data, improving base policy success rates by over 60% on two challenging tasks (book flipping and belt assembly) while outperforming both retraining-from-scratch and finetuning approaches. Through extensive real-world experiments, we provide practical guidance for implementing effective DAgger in real-world robot learning tasks. Result videos are available at: https://compliant-residual-dagger.github.io/ Figure 1: CR-DAgger. To improve a robot manipulation policy, we propose a compliant intervention interface (a) for collecting human correction data, and use this data to update a compliant residual policy (b), and thoroughly study their effects by deploying the updated policy on two contact-rich manipulation tasks in the real world (c). * Equal Contributions 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
succeeds. This process is broadly referred to as Dataset Aggregation (DAgger) [36,23]. However, doing DAgger effectively for real-world robotic problems still faces the following challenges:
How to collect informative human correction data? DAgger is most effective when the correction data is within the original policy's induced state-action distribution [36]. In practice, the common approach is either (1) collecting offline demonstrations that cover the policy's typical failure scenarios [8], or (2) human taking over robot control during policy deployment [37,32]. However, in both cases, the human demonstrator has no access to the original policy's behavior and may deviate excessively from it. Human taking over additionally introduces force discontinuity when they do not instantly reproduce the exact same robot force. This is partially due to the lack of effective correction interfaces that support precise and instantaneous intervention.
How to effectively update the policy with new data? Prior methods for improving a pretrained policy with additional data include (1) retraining the policy from scratch with the aggregated dataset [23], which can be computationally expensive; (2) finetuning the policy with only the additional data [41,17,7], which is sensitive to the quality of the new data [51], and (3) training a residual policy separately on top of the pretrained policy, which is typically done with Reinforcement Learning [2, 51] or Imitation Learning [5], both require a large number of samples.
In this work, we address these questions by proposing an improved system Compliant Residual DAgger (CR-DAgger) consisting of two critical components:
• Compliant Intervention Interface. We propose an on-policy correction system based on kinesthetic teaching to collect delta action without interrupting the current robot policy. Leveraging compliance control, the interface lets humans directly apply force to the robot and feel the magnitude of their instantaneous correction. Unlike take-over corrections, our design allows smooth transition between correction/no correction mode, while providing direct control of correction magnitudes.
• Compliant Residual Policy. Leveraging the force feedback from our Compliant Intervention
Interface, we propose a residual policy formulation that takes in an extra force modality and predicts both residual motions and target forces, which can fully describe the human correction behavior. The Compliant Residual Policy is force-aware, even when the base policy is positiononly. We show that our residual policy formulation learns effective correction strategies using the data collected from our Compliant Intervention Interface.
Together, our system significantly improves the success rate of precise contact-rich robot manipulation tasks using a small amount of additional data. We demonstrate the efficacy of our method on two challenging tasks involving long horizons and sequences of contacts: book flipping and belt assembly. We improve over the base policy success rate by over 60%, while also outperforming retrain-fromscratch and finetuning under the same data budgets. In summary, our contributions are:
• A Compliant Intervention Interface, a system that allows humans to provide accurate, gentle, and smooth corrections in both position and force to a running robot policy without interrupting it.
• A Compliant Residual Pol
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion ModelsRyan Po, Eric Ryan Chan, Changan Chen, Gordon WetzsteinCVPR 2026 · 被引用 18 次
- Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and OpportunitiesChangdae Oh, Seongheon Park, To Eun Kim, Jiatong Li 等ACL 2026 · 被引用 8 次
- What Makes Value Learning Efficient in Residual Reinforcement Learning?Guozheng Ma, Lu Li, Haoyu Wang, Zixuan Liu 等ICML 2026 · 被引用 2 次
- Focus-Then-Contact: Speeding Up Robotic Contact-Rich Task Learning with Affordance-Guided Real-World Residual Reinforcement LearningGuanren Qiao, Ruixiang Ouyang, Sheng Xu, Ruixing Jin 等ICML 2026
它引用的顶会 Paper7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 被引用 326 次
- Flow Matching for Generative ModelingYaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel 等ICLR 2023 · 被引用 87 次
- ASPiRe: Adaptive Skill Priors for Reinforcement LearningMengda Xu, Manuela Veloso, Shuran SongNeurIPS 2022 · 被引用 15 次
- Rapidly Adapting Policies to the Real-World via Simulation-Guided Fine-TuningPatrick Yin, Tyler Westenbroek, Ching-An Cheng, Andrey Kolobov 等ICLR 2025
相关 Paper
- RLIF: Interactive Imitation Learning as Reinforcement LearningJianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma 等ICLR 2024 · 被引用 31 次
- Dexterous Manipulation Transfer via Progressive Kinematic-Dynamic AlignmentWenbin Bai, Qiyu Chen, Xiangbo Lin, Jianwen Li 等AAAI 2026
- The Ingredients of Real World Robotic Reinforcement LearningHenry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah 等ICLR 2020 · 被引用 202 次
- Residual Force Control for Agile Human Behavior Imitation and Extended Motion SynthesisYe Yuan, Kris KitaniNeurIPS 2020 · 被引用 105 次
- DexFlyWheel: A Scalable and Self-improving Data Generation Framework for Dexterous ManipulationKefei Zhu, Fengshuo Bai, YuanHao Xiang, Yishuai Cai 等NeurIPS 2025 · 被引用 8 次
