Lune

NeurIPS2021Top-tier venue

Towards mental time travel: a hierarchical memory for reinforcement learning agents

Andrew K. Lampinen, Stephanie C. Y. Chan, Andrea Banino, Felix Hill

2021Year
63Citations
14Top-tier citations

Abstract

Reinforcement learning agents often forget details of the past, especially after delays or distractor tasks. Agents with common memory architectures struggle to recall and integrate across multiple timesteps of a past event, or even to recall the details of a single timestep that is followed by distractor tasks. To address these limitations, we propose a Hierarchical Chunk Attention Memory (HCAM), which helps agents to remember the past in detail. HCAM stores memories by dividing the past into chunks, and recalls by first performing high-level attention over coarse summaries of the chunks, and then performing detailed attention within only the most relevant chunks. An agent with HCAM can therefore "mentally time-travel"remember past events in detail without attending to all intervening events. We show that agents with HCAM substantially outperform agents with other memory architectures at tasks requiring long-term recall, retention, or reasoning over memory. These include recalling where an object is hidden in a 3D environment, rapidly learning to navigate efficiently in a new neighborhood, and rapidly learning and retaining new object names. Agents with HCAM can extrapolate to task sequences an order of magnitude longer than they were trained on, and can even generalize zero-shot from a meta-learning setting to maintaining knowledge across episodes. HCAM improves agent sample efficiency, generalization, and generality (by solving tasks that previously required specialized architectures). Our work is a step towards agents that can learn, interact, and adapt in complex and temporallyextended environments. [44]-specifically, a gated version of TransformerXL [9] . However, using TransformerXL memories on challenging memory tasks still requires auxiliary self-supervision [19] , and might even benefit from new unsupervised learning mechanisms [3] . Furthermore, even in supervised tasks Transformers can struggle to recall details from long sequences [54] . In the next section, we propose a new hierarchical attention memory architecture for RL agents that can help overcome these challenges. 2 Hierarchical Chunk Attention Memory . . .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7a9ff753-241f-49e3-a7c9-fbc014d5e3c8

Cited by top-tier papers14

Ask how each one uses it

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines