PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic Manipulation
Runze Liu, Yali Du, Fengshuo Bai, Jiafei Lyu, Xiu Li
摘要
In preference-based Reinforcement Learning (RL), obtaining a large number of preference labels are both time-consuming and costly. Furthermore, the queried human preferences cannot be utilized for the new tasks. In this paper, we propose Zero-shot Cross-task Preference Alignment and Robust Reward Learning (PEARL), which learns policies from cross-task preference transfer without any human labels of the target task. Our contributions include two novel components that facilitate the transfer and learning process. The first is Cross-task Preference Alignment (CPA), which transfers the preferences between tasks via optimal transport. The key idea of CPA is to use Gromov-Wasserstein distance to align the trajectories between tasks, and the solved optimal transport matrix serves as the correspondence between trajectories. The target task preferences are computed as the weighted sum of source task preference labels with the correspondence as weights. Moreover, to ensure robust learning from these transferred labels, we introduce Robust Reward Learning (RRL), which considers both reward mean and uncertainty by modeling rewards as Gaussian distributions. Empirical results on robotic manipulation tasks from Meta-World and Robomimic demonstrate that our method is capable of transferring preference labels across tasks accurately and then learns well-behaved policies. Notably, our approach significantly exceeds existing methods when there are few human preferences. The code and videos of our method are available at: https://sites.google. com/view/pearl-preference .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative ReasoningJian Zhao, Runze Liu, Kaiyan Zhang, Zhimu Zhou 等AAAI 2026 · 被引用 68 次
- SARM: Stage-Aware Reward Modeling for Long Horizon Robot ManipulationQianzhong Chen, Justin Yu, Mac Schwager, Pieter Abbeel 等ICLR 2026 · 被引用 50 次
- RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted BehaviorsFengshuo Bai, Runze Liu, Yali Du, Ying Wen 等AAAI 2025 · 被引用 15 次
- RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation SkillsChunru Lin, Haotian Yuan, Yian Wang, Xiaowen Qiu 等NeurIPS 2025 · 被引用 10 次
- Video-Based Optimal Transport for Feedback-Efficient Offline Preference-Based Reinforcement LearningMinh-Tung Luu, Hwanhee Kim, Younghwan Lee, Chang D. YooICML 2026 · 被引用 1 次
它引用的顶会 Paper17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- Graph Optimal Transport for Cross-Domain AlignmentLiqun Chen, Zhe Gan, Yu Cheng, Linjie Li 等ICML 2020 · 被引用 193 次
- Robust Person Re-Identification by Modelling Feature UncertaintyTianyuan Yu, Da Li, Yongxin Yang, Timothy M. Hospedales 等ICCV 2019 · 被引用 148 次
相关 Paper
- What Matters to You? Towards Visual Representation Alignment for Robot LearningThomas Tian, Chenfeng Xu, Masayoshi Tomizuka, Jitendra Malik 等ICLR 2024 · 被引用 17 次
- PAWS: Preference Learning with Advantage-Weighted SegmentsAleksandar Taranovic, Onur Celik, Niklas Freymuth, Ge Li 等ICML 2026
- RIME: Robust Preference-based Reinforcement Learning with Noisy PreferencesJie Cheng, Gang Xiong, Xingyuan Dai, Qinghai Miao 等ICML 2024 · 被引用 42 次
- Alleviating Shifted Distribution in Human Preference Alignment through Meta-LearningShihan Dou, Yan Liu, Enyu Zhou, Songyang Gao 等AAAI 2025 · 被引用 2 次
- Meta-Reinforcement Learning with Adaptation from Human Feedback via Preference-Order-Preserving Task EmbeddingSiyuan Xu, Minghui ZhuICML 2025
