REvolve: Reward Evolution with Large Language Models using Human Feedback
Rishi Hazra, Alkis Sygkounas, Andreas Persson, Amy Loutfi, Pedro Zuidberg Dos Martires
摘要
Designing effective reward functions is crucial to training reinforcement learning (RL) algorithms. However, this design is non-trivial, even for domain experts, due to the subjective nature of certain tasks that are hard to quantify explicitly. In recent works, large language models (LLMs) have been used for reward generation from natural language task descriptions, leveraging their extensive instruction tuning and commonsense understanding of human behavior. In this work, we hypothesize that LLMs, guided by human feedback, can be used to formulate reward functions that reflect human implicit knowledge. We study this in three challenging settings -- autonomous driving, humanoid locomotion, and dexterous manipulation -- wherein notions of ``good" behavior are tacit and hard to quantify. To this end, we introduce REvolve, a truly evolutionary framework that uses LLMs for reward design in RL. REvolve generates and refines reward functions by utilizing human feedback to guide the evolution process, effectively translating implicit human knowledge into explicit reward functions for training (deep) RL agents. Experimentally, we demonstrate that agents trained on REvolve-designed rewards outperform other state-of-the-art baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-benchEdan Toledo, Karen Hambardzumyan, Martin Josifoski, Rishi Hazra 等NeurIPS 2025 · 被引用 71 次
- Partition to Evolve: Niching-enhanced Evolution with LLMs for Automated Algorithm DiscoveryQinglong Hu, Qingfu ZhangNeurIPS 2025 · 被引用 13 次
- RF-Agent: Automated Reward Function Design via Language Agent Tree SearchNing Gao, Xiuhui Zhang, Xingyu Jiang, Mukang You 等NeurIPS 2025 · 被引用 8 次
- Multimodal LLM-assisted Evolutionary Search for Programmatic Control PoliciesQinglong Hu, Tong Xialiang, Mingxuan Yuan, Fei Liu 等ICLR 2026 · 被引用 7 次
- Progress Reward Model for Reinforcement Learning via Large Language ModelsXiuhui Zhang, Ning Gao, Xingyu Jiang, Yihui Chen 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper19
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
相关 Paper
- R*: Efficient Reward Design via Reward Structure Evolution and Parameter Alignment Optimization with Large Language ModelsPengyi Li, Jianye Hao, Hongyao Tang, Yifu Yuan 等ICML 2025
- Efficient Language-instructed Skill Acquisition via Reward-Policy Co-EvolutionChangxin Huang, Yanbin Chang, Junfan Lin, Junyang Liang 等AAAI 2025 · 被引用 1 次
- Text2Reward: Reward Shaping with Language Models for Reinforcement LearningTianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu 等ICLR 2024 · 被引用 142 次
- Learning Reward for Robot Skills Using Large Language Models via Self-AlignmentYuwei Zeng, Yao Mu, Lin ShaoICML 2024 · 被引用 26 次
- Eureka: Human-Level Reward Design via Coding Large Language ModelsYecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang 等ICLR 2024 · 被引用 582 次
