Dual-Scale World Memory for LLM Agents towards Hard-Exploration Problems
Minsoo Kim, Seung-won Hwang
Abstract
LLM-based agents have seen promising advances, yet are still limited in hard-exploration tasks which require agents to perform sustained exploration under sparse feedback. We present GLoW, a novel approach leveraging a dual-scale textual world memory, maintaining a trajectory frontier of high-value discoveries at the global scale, while learning from local trial-and-error in exploration through a Multi-path Advantage Reflection mechanism which infers advantage-based progress signals to guide exploration. To evaluate our framework for hard-exploration, we tackle the Jericho benchmark suite of text-based games, where GLoW achieves a new state-of-the-art performance for LLM-based approaches. Compared to state-of-the-art RL-based methods, our approach achieves comparable performance while requiring 100-800× fewer environment interactions. When scaled to stronger LLMs, GLoW surpasses all prior methods on 4 out of 6 difficult and extreme Jericho games.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 39c84ad0-4f58-44cc-82e2-3e307fdc72f7Builds on16
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- Interactive Fiction Games: A Colossal AdventureMatthew J. Hausknecht, Prithviraj Ammanabrolu, Marc-Alexandre Côté, Xingdi YuanAAAI 2020 · 242 citations
- Graph Constrained Reinforcement Learning for Natural Language Action SpacesPrithviraj Ammanabrolu, Matthew J. HausknechtICLR 2020 · 138 citations
- WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the EnvironmentHao Tang, Darren Key, Kevin EllisNeurIPS 2024 · 123 citations
Related papers
- Keep CALM and Explore: Language Models for Action Generation in Text-based GamesShunyu Yao, Rohan Rao, Matthew J. Hausknecht, Karthik NarasimhanEMNLP 2020 · 67 citations
- Monte Carlo Planning with Large Language Model for Text-Based Game AgentsZijing Shi, Meng Fang, Ling ChenICLR 2025
- Algorithmic Improvements for Deep Reinforcement Learning Applied to Interactive FictionVishal Jain, William Fedus, Hugo Larochelle, Doina Precup et al.AAAI 2020 · 32 citations
- Multi-Stage Episodic Control for Strategic Exploration in Text GamesJens Tuyls, Shunyu Yao, Sham M. Kakade, Karthik NarasimhanICLR 2022 · 30 citations
- SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic StatesZhenliang Zhang, Wenqing Wang, Yong Hu, Yaming Yang et al.ICML 2026
