Compliant Residual DAgger: Improving Real-World Contact-Rich Manipulation with Human Corrections
Xiaomeng Xu, Yifan Hou, Zeyi Liu, Shuran Song
Abstract
We address key challenges in Dataset Aggregation (DAgger) for real-world contactrich manipulation: how to collect informative human correction data and how to effectively update policies with this new data. We introduce Compliant Residual DAgger (CR-DAgger), which contains two novel components: 1) a Compliant Intervention Interface that leverages compliance control, allowing humans to provide gentle, accurate delta action corrections without interrupting the ongoing robot policy execution; and 2) a Compliant Residual Policy formulation that learns from human corrections while incorporating force feedback and force control.
Our system significantly enhances performance on precise contact-rich manipulation tasks using minimal correction data, improving base policy success rates by over 60% on two challenging tasks (book flipping and belt assembly) while outperforming both retraining-from-scratch and finetuning approaches. Through extensive real-world experiments, we provide practical guidance for implementing effective DAgger in real-world robot learning tasks. Result videos are available at: https://compliant-residual-dagger.github.io/ Figure 1: CR-DAgger. To improve a robot manipulation policy, we propose a compliant intervention interface (a) for collecting human correction data, and use this data to update a compliant residual policy (b), and thoroughly study their effects by deploying the updated policy on two contact-rich manipulation tasks in the real world (c). * Equal Contributions 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
succeeds. This process is broadly referred to as Dataset Aggregation (DAgger) [36,23]. However, doing DAgger effectively for real-world robotic problems still faces the following challenges:
How to collect informative human correction data? DAgger is most effective when the correction data is within the original policy's induced state-action distribution [36]. In practice, the common approach is either (1) collecting offline demonstrations that cover the policy's typical failure scenarios [8], or (2) human taking over robot control during policy deployment [37,32]. However, in both cases, the human demonstrator has no access to the original policy's behavior and may deviate excessively from it. Human taking over additionally introduces force discontinuity when they do not instantly reproduce the exact same robot force. This is partially due to the lack of effective correction interfaces that support precise and instantaneous intervention.
How to effectively update the policy with new data? Prior methods for improving a pretrained policy with additional data include (1) retraining the policy from scratch with the aggregated dataset [23], which can be computationally expensive; (2) finetuning the policy with only the additional data [41,17,7], which is sensitive to the quality of the new data [51], and (3) training a residual policy separately on top of the pretrained policy, which is typically done with Reinforcement Learning [2, 51] or Imitation Learning [5], both require a large number of samples.
In this work, we address these questions by proposing an improved system Compliant Residual DAgger (CR-DAgger) consisting of two critical components:
• Compliant Intervention Interface. We propose an on-policy correction system based on kinesthetic teaching to collect delta action without interrupting the current robot policy. Leveraging compliance control, the interface lets humans directly apply force to the robot and feel the magnitude of their instantaneous correction. Unlike take-over corrections, our design allows smooth transition between correction/no correction mode, while providing direct control of correction magnitudes.
• Compliant Residual Policy. Leveraging the force feedback from our Compliant Intervention
Interface, we propose a residual policy formulation that takes in an extra force modality and predicts both residual motions and target forces, which can fully describe the human correction behavior. The Compliant Residual Policy is force-aware, even when the base policy is positiononly. We show that our residual policy formulation learns effective correction strategies using the data collected from our Compliant Intervention Interface.
Together, our system significantly improves the success rate of precise contact-rich robot manipulation tasks using a small amount of additional data. We demonstrate the efficacy of our method on two challenging tasks involving long horizons and sequences of contacts: book flipping and belt assembly. We improve over the base policy success rate by over 60%, while also outperforming retrain-fromscratch and finetuning under the same data budgets. In summary, our contributions are:
• A Compliant Intervention Interface, a system that allows humans to provide accurate, gentle, and smooth corrections in both position and force to a running robot policy without interrupting it.
• A Compliant Residual Pol
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e142e22-bf74-4214-ba42-2a37b5dd02afCited by top-tier papers4
- BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion ModelsRyan Po, Eric Ryan Chan, Changan Chen, Gordon WetzsteinCVPR 2026 · 18 citations
- Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and OpportunitiesChangdae Oh, Seongheon Park, To Eun Kim, Jiatong Li et al.ACL 2026 · 8 citations
- What Makes Value Learning Efficient in Residual Reinforcement Learning?Guozheng Ma, Lu Li, Haoyu Wang, Zixuan Liu et al.ICML 2026 · 2 citations
- Focus-Then-Contact: Speeding Up Robotic Contact-Rich Task Learning with Affordance-Guided Real-World Residual Reinforcement LearningGuanren Qiao, Ruixiang Ouyang, Sheng Xu, Ruixing Jin et al.ICML 2026
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 326 citations
- Flow Matching for Generative ModelingYaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel et al.ICLR 2023 · 87 citations
- ASPiRe: Adaptive Skill Priors for Reinforcement LearningMengda Xu, Manuela Veloso, Shuran SongNeurIPS 2022 · 15 citations
- Rapidly Adapting Policies to the Real-World via Simulation-Guided Fine-TuningPatrick Yin, Tyler Westenbroek, Ching-An Cheng, Andrey Kolobov et al.ICLR 2025
Related papers
- RLIF: Interactive Imitation Learning as Reinforcement LearningJianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma et al.ICLR 2024 · 31 citations
- Dexterous Manipulation Transfer via Progressive Kinematic-Dynamic AlignmentWenbin Bai, Qiyu Chen, Xiangbo Lin, Jianwen Li et al.AAAI 2026
- The Ingredients of Real World Robotic Reinforcement LearningHenry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah et al.ICLR 2020 · 202 citations
- Residual Force Control for Agile Human Behavior Imitation and Extended Motion SynthesisYe Yuan, Kris KitaniNeurIPS 2020 · 105 citations
- DexFlyWheel: A Scalable and Self-improving Data Generation Framework for Dexterous ManipulationKefei Zhu, Fengshuo Bai, YuanHao Xiang, Yishuai Cai et al.NeurIPS 2025 · 8 citations
