ROSETTA: Constructing Code-Based Reward from Unconstrained Language Preference
Sanjana Srivastava, Kangrui Wang, Yung-Chieh Chan, Tianyuan Dai, Manling Li, Ruohan Zhang, Mengdi Xu, Jiajun Wu, Li Fei-Fei
摘要
Intelligent embodied agents not only need to accomplish preset tasks, but also learn to align with individual human needs and preferences. Extracting reward signals from human language preferences allows an embodied agent to adapt through reinforcement learning. However, human language preferences are unconstrained, diverse, and dynamic, making constructing learnable reward from them a major challenge. We present ROSETTA, a framework that uses foundation models to ground and disambiguate unconstrained natural language preference, construct multi-stage reward functions, and implement them with code generation. Unlike prior works requiring extensive offline training to get general reward models or fine-grained correction on a single task, ROSETTA allows agents to adapt online to preference that evolves and is diverse in language and content. We test ROSETTA on both short-horizon and long-horizon manipulation tasks and conduct extensive human evaluation, finding that ROSETTA outperforms SOTA baselines and achieves 87% average success rate and 86% human satisfaction across 116 preferences. Push the ball onto the target. We validate ROSETTA iteratively: multiple steps of taking a language preference, generating a reward, training an agent, evaluating, and taking a new preference. We evaluate on 35 sequences of two to four preferences each, total 116 preferences, in five task-agnostic manipulation environments (Fig. 1 ). ROSETTA successfully interprets ambiguous language, adapts to unseen preferences even after four interaction steps, and produces semantically matched and optimizable rewards that result in aligned
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 被引用 1,539 次
- Eureka: Human-Level Reward Design via Coding Large Language ModelsYecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang 等ICLR 2024 · 被引用 582 次
- FoundationPose: Unified 6D Pose Estimation and Tracking of Novel ObjectsBowen Wen, Wei Yang, Jan Kautz, Stan BirchfieldCVPR 2024 · 被引用 215 次
- Text2Reward: Reward Shaping with Language Models for Reinforcement LearningTianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu 等ICLR 2024 · 被引用 142 次
- RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model FeedbackYufei Wang, Zhanyi Sun, Jesse Zhang, Zhou Xian 等ICML 2024 · 被引用 135 次
相关 Paper
- VLP: Vision-Language Preference Learning for Embodied ManipulationRunze Liu, Chenjia Bai, Jiafei Lyu, Shengjie Sun 等EMNLP 2025 · 被引用 1 次
- Master Skill Learning with Policy-Grounded Synergy of LLM-based Reward Shaping and ExploringYanbin Chang, Junfan Lin, Jie Jiang, Runhao Zeng 等ICLR 2026
- A Regret Minimization Framework on Preference Learning in Large Language ModelsSuhwan Kim, Taehyun Cho, Youngsoo Jang, Geon-Hyeong Kim 等ICML 2026
- Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual AlignmentZhaofeng Wu, Ananth Balashankar, Yoon Kim, Jacob Eisenstein 等EMNLP 2024 · 被引用 29 次
- Personalizing Reinforcement Learning from Human Feedback with Variational Preference LearningSriyash Poddar, Yanming Wan, Hamish Ivison, Abhishek Gupta 等NeurIPS 2024 · 被引用 188 次
