History Compression via Language Models in Reinforcement Learning
Fabian Paischer, Thomas Adler, Vihang Patil, Angela Bitto-Nemling, Markus Holzleitner, Sebastian Lehner, Hamid Eghbal-Zadeh, Sepp Hochreiter
Abstract
In a partially observable Markov decision process (POMDP), an agent typically uses a representation of the past to approximate the underlying MDP. We propose to utilize a frozen Pretrained Language Transformer (PLT) for history representation and compression to improve sample efficiency. To avoid training of the Transformer, we introduce FrozenHopfield, which automatically associates observations with pretrained token embeddings. To form these associations, a modern Hopfield network stores these token embeddings, which are retrieved by queries that are obtained by a random but fixed projection of observations. Our new method, HELM, enables actor-critic network architectures that contain a pretrained language Transformer for history representation as a memory module. Since a representation of the past need not be learned, HELM is much more sample efficient than competitors. On Minigrid and Procgen environments HELM achieves new state-of-the-art results. Our code is available at https://github.com/ml-jku/helm .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIPAndreas Fürst, Elisabeth Rumetshofer, Johannes Lehner, Viet T. Tran et al.NeurIPS 2022 · 131 citations
- Conformal Prediction for Time Series with Modern Hopfield NetworksAndreas Auer, Martin Gauch, Daniel Klotz, Sepp HochreiterNeurIPS 2023 · 70 citations
- On Sparse Modern Hopfield ModelJerry Yao-Chieh Hu, Donglin Yang, Dennis Wu, Chenwei Xu et al.NeurIPS 2023 · 52 citations
- On Computational Limits of Modern Hopfield Models: A Fine-Grained Complexity AnalysisJerry Yao-Chieh Hu, Thomas Lin, Zhao Song, Han LiuICML 2024 · 47 citations
- Align-RUDDER: Learning From Few Demonstrations by Reward RedistributionVihang Patil, Markus Hofmarcher, Marius-Constantin Dinu, Matthias Dorfer et al.ICML 2022 · 46 citations
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
Related papers
- From Tokens to Latent States: Leveraging Pre-trained Language Models for Improving Partially Observable Reinforcement LearningMeiju Li, Ruixiang Sun, Xin Li, Mingzhong WangAAAI 2026
- HiT-MDP: Learning the SMDP option framework on MDPs with Hidden Temporal EmbeddingsChang Li, Dongjin Song, Dacheng TaoICLR 2023
- RePreM: Representation Pre-training with Masked Model for Reinforcement LearningYuanying Cai, Chuheng Zhang, Wei Shen, Xuyun Zhang et al.AAAI 2023 · 7 citations
- Semantic HELM: A Human-Readable Memory for Reinforcement LearningFabian Paischer, Thomas Adler, Markus Hofmarcher, Sepp HochreiterNeurIPS 2023 · 21 citations
- Frozen Pretrained Transformers as Universal Computation EnginesKevin Lu, Aditya Grover, Pieter Abbeel, Igor MordatchAAAI 2022 · 133 citations
