Lune

NeurIPS2025顶会

Compliant Residual DAgger: Improving Real-World Contact-Rich Manipulation with Human Corrections

Xiaomeng Xu, Yifan Hou, Zeyi Liu, Shuran Song

2025年份
57被引次数
4顶会引用

摘要

We address key challenges in Dataset Aggregation (DAgger) for real-world contactrich manipulation: how to collect informative human correction data and how to effectively update policies with this new data. We introduce Compliant Residual DAgger (CR-DAgger), which contains two novel components: 1) a Compliant Intervention Interface that leverages compliance control, allowing humans to provide gentle, accurate delta action corrections without interrupting the ongoing robot policy execution; and 2) a Compliant Residual Policy formulation that learns from human corrections while incorporating force feedback and force control.

Our system significantly enhances performance on precise contact-rich manipulation tasks using minimal correction data, improving base policy success rates by over 60% on two challenging tasks (book flipping and belt assembly) while outperforming both retraining-from-scratch and finetuning approaches. Through extensive real-world experiments, we provide practical guidance for implementing effective DAgger in real-world robot learning tasks. Result videos are available at: https://compliant-residual-dagger.github.io/ Figure 1: CR-DAgger. To improve a robot manipulation policy, we propose a compliant intervention interface (a) for collecting human correction data, and use this data to update a compliant residual policy (b), and thoroughly study their effects by deploying the updated policy on two contact-rich manipulation tasks in the real world (c). * Equal Contributions 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

succeeds. This process is broadly referred to as Dataset Aggregation (DAgger) [36,23]. However, doing DAgger effectively for real-world robotic problems still faces the following challenges:

How to collect informative human correction data? DAgger is most effective when the correction data is within the original policy's induced state-action distribution [36]. In practice, the common approach is either (1) collecting offline demonstrations that cover the policy's typical failure scenarios [8], or (2) human taking over robot control during policy deployment [37,32]. However, in both cases, the human demonstrator has no access to the original policy's behavior and may deviate excessively from it. Human taking over additionally introduces force discontinuity when they do not instantly reproduce the exact same robot force. This is partially due to the lack of effective correction interfaces that support precise and instantaneous intervention.

How to effectively update the policy with new data? Prior methods for improving a pretrained policy with additional data include (1) retraining the policy from scratch with the aggregated dataset [23], which can be computationally expensive; (2) finetuning the policy with only the additional data [41,17,7], which is sensitive to the quality of the new data [51], and (3) training a residual policy separately on top of the pretrained policy, which is typically done with Reinforcement Learning [2, 51] or Imitation Learning [5], both require a large number of samples.

In this work, we address these questions by proposing an improved system Compliant Residual DAgger (CR-DAgger) consisting of two critical components:

• Compliant Intervention Interface. We propose an on-policy correction system based on kinesthetic teaching to collect delta action without interrupting the current robot policy. Leveraging compliance control, the interface lets humans directly apply force to the robot and feel the magnitude of their instantaneous correction. Unlike take-over corrections, our design allows smooth transition between correction/no correction mode, while providing direct control of correction magnitudes.

• Compliant Residual Policy. Leveraging the force feedback from our Compliant Intervention

Interface, we propose a residual policy formulation that takes in an extra force modality and predicts both residual motions and target forces, which can fully describe the human correction behavior. The Compliant Residual Policy is force-aware, even when the base policy is positiononly. We show that our residual policy formulation learns effective correction strategies using the data collected from our Compliant Intervention Interface.

Together, our system significantly improves the success rate of precise contact-rich robot manipulation tasks using a small amount of additional data. We demonstrate the efficacy of our method on two challenging tasks involving long horizons and sequences of contacts: book flipping and belt assembly. We improve over the base policy success rate by over 60%, while also outperforming retrain-fromscratch and finetuning under the same data budgets. In summary, our contributions are:

• A Compliant Intervention Interface, a system that allows humans to provide accurate, gentle, and smooth corrections in both position and force to a running robot policy without interrupting it.

• A Compliant Residual Pol

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper4

问问它们各自怎么用它

它引用的顶会 Paper7

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖