Promptable Behaviors: Personalizing Multi-Objective Rewards from Human Preferences
Minyoung Hwang, Luca Weihs, Chanwoo Park, Kimin Lee, Aniruddha Kembhavi, Kiana Ehsani
摘要
Customizing robotic behaviors to be aligned with diverse human preferences is an underexplored challenge in the field of embodied AI. In this paper, we present Promptable Behaviors, a novel framework that facilitates efficient personalization of robotic agents to diverse human preferences in complex environments. We use multi-objective reinforcement learning to train a single policy adaptable to a broad spectrum of preferences. We introduce three distinct methods to infer human preferences by leveraging different types of interactions: (1) human demonstrations, (2) preference feedback on trajectory comparisons, and (3) language instructions. We evaluate the proposed method in personalized object-goal navigation and flee navigation tasks in ProcTHOR [18] and RoboTHOR [17] , demonstrating the ability to prompt agent behaviors to satisfy human preferences in various scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Panacea: Pareto Alignment via Preference Adaptation for LLMsYifan Zhong, Chengdong Ma, Xiaoyuan Zhang, Ziran Yang 等NeurIPS 2024 · 被引用 89 次
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu 等ACL 2025 · 被引用 45 次
- PSG-Nav: Probabilistic Scene Graph Navigation via Multiverse Decision MakingRufeng Chen, Yue Chang, Xiaqiang Tang, Hechang Chen 等ICML 2026 · 被引用 3 次
- Comparison-based Active Preference Learning for Multi-dimensional PersonalizationMinhyeon Oh, Seungjoon Lee, Jungseul OkACL 2025 · 被引用 1 次
- MERIT Feedback Elicits Better Bargaining in LLM NegotiatorsJihwan Oh, Murad Aghazada, Yooju Shin, Se-Young Yun 等ACL 2026 · 被引用 1 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee 等ICLR 2020 · 被引用 608 次
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus 等ICML 2020 · 被引用 210 次
相关 Paper
- AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion ModelZibin Dong, Yifu Yuan, Jianye Hao, Fei Ni 等ICLR 2024 · 被引用 44 次
- Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement LearningTianmeng Hu, Biao Luo, Ke LiICML 2026 · 被引用 3 次
- Relative Behavioral Attributes: Filling the Gap between Symbolic Goal Specification and Reward Learning from Human PreferencesLin Guan, Karthik Valmeekam, Subbarao KambhampatiICLR 2023
- An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement LearningQian Lin, Zongkai Liu, Danying Mo, Chao YuNeurIPS 2024 · 被引用 8 次
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert 等ICML 2020 · 被引用 93 次
