Reinforcement Learning with Simple Sequence Priors
Tankred Saanum, Noémi Élteto, Peter Dayan, Marcel Binz, Eric Schulz
Abstract
Everything else being equal, simpler models should be preferred over more complex ones. In reinforcement learning (RL), simplicity is typically quantified on an actionby-action basis -but this timescale ignores temporal regularities, like repetitions, often present in sequential strategies. We therefore propose an RL algorithm that learns to solve tasks with sequences of actions that are compressible. We explore two possible sources of simple action sequences: Sequences that can be learned by autoregressive models, and sequences that are compressible with off-the-shelf data compression algorithms. Distilling these preferences into sequence priors, we derive a novel information-theoretic objective that incentivizes agents to learn policies that maximize rewards while conforming to these priors. We show that the resulting RL algorithm leads to faster learning, and attains higher returns than state-of-the-art model-free approaches in a series of continuous control tasks from the DeepMind Control Suite. These priors also produce a powerful informationregularized agent that is robust to noisy observations and can perform open-loop control. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37bbc5b7-5583-4b89-8404-5a5faeecf879Cited by top-tier papers7
- Simplifying Latent Dynamics with Softly State-Invariant World ModelsTankred Saanum, Peter Dayan, Eric SchulzNeurIPS 2024 · 14 citations
- Evaluating alignment between humans and neural network representations in image-based learning tasksCan Demircan, Tankred Saanum, Leonardo Pettini, Marcel Binz et al.NeurIPS 2024 · 11 citations
- Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement LearningYounggyo Seo, Pieter AbbeelNeurIPS 2025 · 11 citations
- Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse RewardJiarui Yang, Bin Zhu, Jingjing Chen, Yu-Gang JiangAAAI 2026 · 3 citations
- Sparse Autoencoders Reveal Temporal Difference Learning in Large Language ModelsCan Demircan, Tankred Saanum, Akshay Kumar Jagadish, Marcel Binz et al.ICLR 2025 · 1 citation
Builds on9
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Improving Sample Efficiency in Model-Free Reinforcement Learning from ImagesDenis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos et al.AAAI 2021 · 506 citations
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 244 citations
- Learning Invariant Representations for Reinforcement Learning without ReconstructionAmy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal et al.ICLR 2021 · 77 citations
Related papers
- Robust Predictable ControlBen Eysenbach, Ruslan Salakhutdinov, Sergey LevineNeurIPS 2021 · 53 citations
- Temporal Predictive Coding For Model-Based Planning In Latent SpaceTung D. Nguyen, Rui Shu, Tuan Pham, Hung Bui et al.ICML 2021 · 65 citations
- Efficient Unsupervised Sentence Compression by Fine-tuning Transformers with Reinforcement LearningDemian Gholipour Ghalandari, Chris Hokamp, Georgiana IfrimACL 2022
- Exploration via Planning for Information about the Optimal TrajectoryViraj Mehta, Ian Char, Joseph Abbate, Rory Conlin et al.NeurIPS 2022 · 12 citations
- Overcoming Slow Decision Frequencies in Continuous Control: Model-Based Sequence Reinforcement Learning for Model-Free ControlDevdhar Patel, Hava T. SiegelmannICLR 2025
