Learning Dynamic Belief Graphs to Generalize on Text-Based Games
Ashutosh Adhikari, Xingdi Yuan, Marc-Alexandre Côté, Mikulas Zelinka, Marc-Antoine Rondeau, Romain Laroche, Pascal Poupart, Jian Tang, Adam Trischler, William L. Hamilton
Abstract
Playing text-based games requires skills in processing natural language and sequential decision making. Achieving human-level performance on text-based games remains an open challenge, and prior research has largely relied on hand-crafted structured representations and heuristics. In this work, we investigate how an agent can plan and generalize in text-based games using graph-structured representations learned end-to-end from raw text. We propose a novel graph-aided transformer agent (GATA) that infers and updates latent belief graphs during planning to enable effective action selection by capturing the underlying game dynamics. GATA is trained using a combination of reinforcement and self-supervised learning. Our work demonstrates that the learned graph-based representations help agents converge to better policies than their text-only counterparts and facilitate effective generalization across game configurations. Experiments on 500+ unique games from the TextWorld suite show that our best agent outperforms text-based baselines by an average of 24.2%. * Equal contribution. 2 We challenge readers to solve this representative game: https://aka.ms/textworld-tryit .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd32db85-efb8-4014-b523-c0ceb8e045aaCited by top-tier papers18
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
- Can Language Models Solve Graph Problems in Natural Language?Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan et al.NeurIPS 2023 · 420 citations
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli BenchmarkAlexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li et al.ICML 2023 · 200 citations
- Large Language Models Are Neurosymbolic ReasonersMeng Fang, Shilong Deng, Yudi Zhang, Zijing Shi et al.AAAI 2024 · 53 citations
- Learning Knowledge Graph-based World Models of Textual EnvironmentsPrithviraj Ammanabrolu, Mark O. RiedlNeurIPS 2021 · 43 citations
Builds on7
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 322 citations
- Interactive Fiction Games: A Colossal AdventureMatthew J. Hausknecht, Prithviraj Ammanabrolu, Marc-Alexandre Côté, Xingdi YuanAAAI 2020 · 242 citations
- Graph Constrained Reinforcement Learning for Natural Language Action SpacesPrithviraj Ammanabrolu, Matthew J. HausknechtICLR 2020 · 138 citations
Related papers
- Deep Reinforcement Learning with Stacked Hierarchical Attention for Text-based GamesYunqiu Xu, Meng Fang, Ling Chen, Yali Du et al.NeurIPS 2020 · 48 citations
- Eye of the Beholder: Improved Relation Generalization for Text-Based Reinforcement Learning AgentsKeerthiram Murugesan, Subhajit Chaudhury, Kartik TalamadupulaAAAI 2022 · 5 citations
- Learning Symbolic Rules over Abstract Meaning Representations for Textual Reinforcement LearningSubhajit Chaudhury, Sarathkrishna Swaminathan, Daiki Kimura, Prithviraj Sen et al.ACL 2023 · 4 citations
- Theory of Mind for Multi-Agent Collaboration via Large Language ModelsHuao Li, Yu Quan Chong, Simon Stepputtis, Joseph Campbell et al.EMNLP 2023 · 57 citations
- Learning Object-Oriented Dynamics for Planning from TextGuiliang Liu, Ashutosh Adhikari, Amir-massoud Farahmand, Pascal PoupartICLR 2022 · 9 citations
