EPO: Hierarchical LLM Agents with Environment Preference Optimization
Qi Zhao, Haotian Fu, Chen Sun, George Konidaris
Abstract
Long-horizon decision-making tasks present significant challenges for LLM-based agents due to the need for extensive planning over multiple steps. In this paper, we propose a hierarchical framework that decomposes complex tasks into manageable subgoals, utilizing separate LLMs for subgoal prediction and lowlevel action generation. To address the challenge of creating training signals for unannotated datasets, we develop a reward model that leverages multimodal environment feedback to automatically generate reward signals. We introduce Environment Preference Optimization (EPO), a novel method that generates preference signals from the environment's feedback and uses them to train LLM-based agents. Extensive experiments on ALFRED demonstrate the state-of-the-art performance of our framework, achieving first place on the ALFRED public leaderboard and showcasing its potential to improve long-horizon decision-making in diverse environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bba8b312-9a96-460d-aee0-7e1bb443e43bCited by top-tier papers9
- Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMsYifan Shen, Yuanzhe Liu, Jingyuan Zhu, Xu Cao et al.NeurIPS 2025 · 41 citations
- World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task PlanningSiyin Wang, Zhaoye Fei, Qinyuan Cheng, Shiduo Zhang et al.ACL 2025 · 16 citations
- Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMsChanghao Li, Yuchen Zhuang, Rushi Qiang, Haotian Sun et al.NeurIPS 2025 · 12 citations
- World-aware Planning Narratives Enhance Large Vision-Language Model PlannerJunhao Shi, Zhaoye Fei, Siyin Wang, Qipeng Guo et al.NeurIPS 2025 · 11 citations
- ReCAPA: Hierarchical Predictive Correction to Mitigate Cascading FailuresXiyin Zeng, Yuyu Sun, Haoyang Li, Shouqiang Liu et al.ICLR 2026 · 2 citations
Builds on20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
Related papers
- Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM AgentsHeyang Gao, Zexu Sun, Erxue Min, Hengyi Cai et al.ICLR 2026 · 5 citations
- Scaling Autonomous Agents via Automatic Reward Modeling And PlanningZhenfang Chen, Delin Chen, Rui Sun, Wenjun Liu et al.ICLR 2025
- Instruction-Augmented Long-Horizon Planning: Embedding Grounding Mechanisms in Embodied Mobile ManipulationFangyuan Wang, Shipeng Lyu, Peng Zhou, Anqing Duan et al.AAAI 2025 · 9 citations
- Episodic Transformer for Vision-and-Language NavigationAlexander Pashevich, Cordelia Schmid, Chen SunICCV 2021 · 228 citations
- BaseReward: A Strong Baseline for Multimodal Reward ModelYiFan Zhang, Haihua Yang, Huanyu Zhang, Yang Shi et al.ICLR 2026 · 16 citations
