EAGER: Asking and Answering Questions for Automatic Reward Shaping in Language-guided RL
Thomas Carta, Pierre-Yves Oudeyer, Olivier Sigaud, Sylvain Lamprier
摘要
Reinforcement learning (RL) in long horizon and sparse reward tasks is notoriously difficult and requires a lot of training steps. A standard solution to speed up the process is to leverage additional reward signals, shaping it to better guide the learning process. In the context of language-conditioned RL, the abstraction and generalisation properties of the language input provide opportunities for more efficient ways of shaping the reward. In this paper, we leverage this idea and propose an automated reward shaping method where the agent extracts auxiliary objectives from the general language goal. These auxiliary objectives use a question generation (QG) and question answering (QA) system: they consist of questions leading the agent to try to reconstruct partial information about the global goal using its own trajectory. When it succeeds, it receives an intrinsic reward proportional to its confidence in its answer. This incentivizes the agent to generate trajectories which unambiguously explain various aspects of the general language goal. Our experimental study shows that this approach, which does not require engineer intervention to design the auxiliary objectives, improves sample efficiency by effectively directing exploration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Reward Design with Language ModelsMinae Kwon, Sang Michael Xie, Kalesha Bullard, Dorsa SadighICLR 2023 · 被引用 21 次
- Semantic HELM: A Human-Readable Memory for Reinforcement LearningFabian Paischer, Thomas Adler, Markus Hofmarcher, Sepp HochreiterNeurIPS 2023 · 被引用 21 次
- Learning with Language-Guided State AbstractionsAndi Peng, Ilia Sucholutsky, Belinda Z. Li, Theodore R. Sumers 等ICLR 2024 · 被引用 20 次
- Offline RL by Reward-Weighted Fine-Tuning for Conversation OptimizationSubhojyoti Mukherjee, Viet Dac Lai, Raghavendra Addanki, Ryan Rossi 等NeurIPS 2025 · 被引用 12 次
- The Bidirectional Process Reward ModelLingyin Zhang, Jun Gao, Xiaoxue Ren, Ziqiang CaoACL 2026 · 被引用 2 次
它引用的顶会 Paper11
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- Episodic Transformer for Vision-and-Language NavigationAlexander Pashevich, Cordelia Schmid, Chen SunICCV 2021 · 被引用 228 次
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 被引用 198 次
- Language as a Cognitive Tool to Imagine Goals in Curiosity Driven ExplorationCédric Colas, Tristan Karch, Nicolas Lair, Jean-Michel Dussoux 等NeurIPS 2020 · 被引用 139 次
相关 Paper
- Reward Shaping for Reinforcement Learning with An Assistant Reward AgentHaozhe Ma, Kuankuan Sima, Thanh Vinh Vo, Di Fu 等ICML 2024 · 被引用 34 次
- Induction of Subgoal Automata for Reinforcement LearningDaniel Furelos-Blanco, Mark Law, Alessandra Russo, Krysia Broda 等AAAI 2020 · 被引用 37 次
- Strategic Planning: A Top-Down Approach to Option GenerationMax Ruiz Luyten, Antonin Berthon, Mihaela van der SchaarICML 2025
- Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample ComplexityAbhishek Gupta, Aldo Pacchiano, Yuexiang Zhai, Sham M. Kakade 等NeurIPS 2022 · 被引用 115 次
- Instructing Goal-Conditioned Reinforcement Learning Agents with Temporal Logic ObjectivesWenjie Qiu, Wensen Mao, He ZhuNeurIPS 2023 · 被引用 44 次
