ROSETTA: Constructing Code-Based Reward from Unconstrained Language Preference
Sanjana Srivastava, Kangrui Wang, Yung-Chieh Chan, Tianyuan Dai, Manling Li, Ruohan Zhang, Mengdi Xu, Jiajun Wu, Li Fei-Fei
Abstract
Intelligent embodied agents not only need to accomplish preset tasks, but also learn to align with individual human needs and preferences. Extracting reward signals from human language preferences allows an embodied agent to adapt through reinforcement learning. However, human language preferences are unconstrained, diverse, and dynamic, making constructing learnable reward from them a major challenge. We present ROSETTA, a framework that uses foundation models to ground and disambiguate unconstrained natural language preference, construct multi-stage reward functions, and implement them with code generation. Unlike prior works requiring extensive offline training to get general reward models or fine-grained correction on a single task, ROSETTA allows agents to adapt online to preference that evolves and is diverse in language and content. We test ROSETTA on both short-horizon and long-horizon manipulation tasks and conduct extensive human evaluation, finding that ROSETTA outperforms SOTA baselines and achieves 87% average success rate and 86% human satisfaction across 116 preferences. Push the ball onto the target. We validate ROSETTA iteratively: multiple steps of taking a language preference, generating a reward, training an agent, evaluating, and taking a new preference. We evaluate on 35 sequences of two to four preferences each, total 116 preferences, in five task-agnostic manipulation environments (Fig. 1 ). ROSETTA successfully interprets ambiguous language, adapts to unseen preferences even after four interaction steps, and produces semantically matched and optimizable rewards that result in aligned
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e5dac69-7f01-48fd-9c74-942f687a9ad4Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- Eureka: Human-Level Reward Design via Coding Large Language ModelsYecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang et al.ICLR 2024 · 582 citations
- FoundationPose: Unified 6D Pose Estimation and Tracking of Novel ObjectsBowen Wen, Wei Yang, Jan Kautz, Stan BirchfieldCVPR 2024 · 215 citations
- Text2Reward: Reward Shaping with Language Models for Reinforcement LearningTianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu et al.ICLR 2024 · 142 citations
- RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model FeedbackYufei Wang, Zhanyi Sun, Jesse Zhang, Zhou Xian et al.ICML 2024 · 135 citations
Related papers
- VLP: Vision-Language Preference Learning for Embodied ManipulationRunze Liu, Chenjia Bai, Jiafei Lyu, Shengjie Sun et al.EMNLP 2025 · 1 citation
- Master Skill Learning with Policy-Grounded Synergy of LLM-based Reward Shaping and ExploringYanbin Chang, Junfan Lin, Jie Jiang, Runhao Zeng et al.ICLR 2026
- A Regret Minimization Framework on Preference Learning in Large Language ModelsSuhwan Kim, Taehyun Cho, Youngsoo Jang, Geon-Hyeong Kim et al.ICML 2026
- Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual AlignmentZhaofeng Wu, Ananth Balashankar, Yoon Kim, Jacob Eisenstein et al.EMNLP 2024 · 29 citations
- Personalizing Reinforcement Learning from Human Feedback with Variational Preference LearningSriyash Poddar, Yanming Wan, Hamish Ivison, Abhishek Gupta et al.NeurIPS 2024 · 188 citations
