Learning Reward for Robot Skills Using Large Language Models via Self-Alignment
Yuwei Zeng, Yao Mu, Lin Shao
Abstract
Learning reward functions remains the bottleneck to equip a robot with a broad repertoire of skills. Large Language Models (LLM) contain valuable task-related knowledge that can potentially aid in the learning of reward functions. However, the proposed reward function can be imprecise, thus ineffective which requires to be further grounded with environment information. We proposed a method to learn rewards more efficiently in the absence of humans. Our approach consists of two components: We first use the LLM to propose features and parameterization of the reward, then update the parameters through an iterative self-alignment process. In particular, the process minimizes the ranking inconsistency between the LLM and the learnt reward functions based on the execution feedback. The method was validated on 9 tasks across 2 simulation environments. It demonstrates a consistent improvement over training efficacy and efficiency, meanwhile consuming significantly fewer GPT tokens compared to the alternative mutation-based method. Project website: https://sites.google.com/view/rewardselfalign .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e553de97-44fc-4797-8456-5a21d1a6b97eCited by top-tier papers9
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi et al.EMNLP 2024 · 119 citations
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation SkillsChunru Lin, Haotian Yuan, Yian Wang, Xiaowen Qiu et al.NeurIPS 2025 · 10 citations
- RF-Agent: Automated Reward Function Design via Language Agent Tree SearchNing Gao, Xiuhui Zhang, Xingyu Jiang, Mukang You et al.NeurIPS 2025 · 8 citations
- ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied TasksPhilip Schroeder, Ondrej Biza, Thomas Weng, Hongyin Luo et al.NeurIPS 2025 · 8 citations
Builds on7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative SimulationYufei Wang, Zhou Xian, Feng Chen, Tsun-Hsuan Wang et al.ICML 2024 · 227 citations
- Tree-Planner: Efficient Close-loop Task Planning with Large Language ModelsMengkang Hu, Yao Mu, Xinmiao Yu, Mingyu Ding et al.ICLR 2024 · 57 citations
Related papers
- R*: Efficient Reward Design via Reward Structure Evolution and Parameter Alignment Optimization with Large Language ModelsPengyi Li, Jianye Hao, Hongyao Tang, Yifu Yuan et al.ICML 2025
- REvolve: Reward Evolution with Large Language Models using Human FeedbackRishi Hazra, Alkis Sygkounas, Andreas Persson, Amy Loutfi et al.ICLR 2025
- Efficient Language-instructed Skill Acquisition via Reward-Policy Co-EvolutionChangxin Huang, Yanbin Chang, Junfan Lin, Junyang Liang et al.AAAI 2025 · 1 citation
- Aligning Large Language Models through Synthetic FeedbackSungdong Kim, Sanghwan Bae, Jamin Shin, Soyoung Kang et al.EMNLP 2023 · 13 citations
- Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language ModelsSomanshu Singla, Zhen Wang, Tianyang Liu, Abdullah Ashfaq et al.EMNLP 2024 · 1 citation
