Lune

NeurIPS2025Top-tier venue

Compliant Residual DAgger: Improving Real-World Contact-Rich Manipulation with Human Corrections

Xiaomeng Xu, Yifan Hou, Zeyi Liu, Shuran Song

2025Year
57Citations
4Top-tier citations

Abstract

We address key challenges in Dataset Aggregation (DAgger) for real-world contactrich manipulation: how to collect informative human correction data and how to effectively update policies with this new data. We introduce Compliant Residual DAgger (CR-DAgger), which contains two novel components: 1) a Compliant Intervention Interface that leverages compliance control, allowing humans to provide gentle, accurate delta action corrections without interrupting the ongoing robot policy execution; and 2) a Compliant Residual Policy formulation that learns from human corrections while incorporating force feedback and force control.

Our system significantly enhances performance on precise contact-rich manipulation tasks using minimal correction data, improving base policy success rates by over 60% on two challenging tasks (book flipping and belt assembly) while outperforming both retraining-from-scratch and finetuning approaches. Through extensive real-world experiments, we provide practical guidance for implementing effective DAgger in real-world robot learning tasks. Result videos are available at: https://compliant-residual-dagger.github.io/ Figure 1: CR-DAgger. To improve a robot manipulation policy, we propose a compliant intervention interface (a) for collecting human correction data, and use this data to update a compliant residual policy (b), and thoroughly study their effects by deploying the updated policy on two contact-rich manipulation tasks in the real world (c). * Equal Contributions 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

succeeds. This process is broadly referred to as Dataset Aggregation (DAgger) [36,23]. However, doing DAgger effectively for real-world robotic problems still faces the following challenges:

How to collect informative human correction data? DAgger is most effective when the correction data is within the original policy's induced state-action distribution [36]. In practice, the common approach is either (1) collecting offline demonstrations that cover the policy's typical failure scenarios [8], or (2) human taking over robot control during policy deployment [37,32]. However, in both cases, the human demonstrator has no access to the original policy's behavior and may deviate excessively from it. Human taking over additionally introduces force discontinuity when they do not instantly reproduce the exact same robot force. This is partially due to the lack of effective correction interfaces that support precise and instantaneous intervention.

How to effectively update the policy with new data? Prior methods for improving a pretrained policy with additional data include (1) retraining the policy from scratch with the aggregated dataset [23], which can be computationally expensive; (2) finetuning the policy with only the additional data [41,17,7], which is sensitive to the quality of the new data [51], and (3) training a residual policy separately on top of the pretrained policy, which is typically done with Reinforcement Learning [2, 51] or Imitation Learning [5], both require a large number of samples.

In this work, we address these questions by proposing an improved system Compliant Residual DAgger (CR-DAgger) consisting of two critical components:

• Compliant Intervention Interface. We propose an on-policy correction system based on kinesthetic teaching to collect delta action without interrupting the current robot policy. Leveraging compliance control, the interface lets humans directly apply force to the robot and feel the magnitude of their instantaneous correction. Unlike take-over corrections, our design allows smooth transition between correction/no correction mode, while providing direct control of correction magnitudes.

• Compliant Residual Policy. Leveraging the force feedback from our Compliant Intervention

Interface, we propose a residual policy formulation that takes in an extra force modality and predicts both residual motions and target forces, which can fully describe the human correction behavior. The Compliant Residual Policy is force-aware, even when the base policy is positiononly. We show that our residual policy formulation learns effective correction strategies using the data collected from our Compliant Intervention Interface.

Together, our system significantly improves the success rate of precise contact-rich robot manipulation tasks using a small amount of additional data. We demonstrate the efficacy of our method on two challenging tasks involving long horizons and sequences of contacts: book flipping and belt assembly. We improve over the base policy success rate by over 60%, while also outperforming retrain-fromscratch and finetuning under the same data budgets. In summary, our contributions are:

• A Compliant Intervention Interface, a system that allows humans to provide accurate, gentle, and smooth corrections in both position and force to a running robot policy without interrupting it.

• A Compliant Residual Pol

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 3e142e22-bf74-4214-ba42-2a37b5dd02af

Cited by top-tier papers4

Ask how each one uses it

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines