Reinforcement Learning with Simple Sequence Priors
Tankred Saanum, Noémi Élteto, Peter Dayan, Marcel Binz, Eric Schulz
摘要
Everything else being equal, simpler models should be preferred over more complex ones. In reinforcement learning (RL), simplicity is typically quantified on an actionby-action basis -but this timescale ignores temporal regularities, like repetitions, often present in sequential strategies. We therefore propose an RL algorithm that learns to solve tasks with sequences of actions that are compressible. We explore two possible sources of simple action sequences: Sequences that can be learned by autoregressive models, and sequences that are compressible with off-the-shelf data compression algorithms. Distilling these preferences into sequence priors, we derive a novel information-theoretic objective that incentivizes agents to learn policies that maximize rewards while conforming to these priors. We show that the resulting RL algorithm leads to faster learning, and attains higher returns than state-of-the-art model-free approaches in a series of continuous control tasks from the DeepMind Control Suite. These priors also produce a powerful informationregularized agent that is robust to noisy observations and can perform open-loop control. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Simplifying Latent Dynamics with Softly State-Invariant World ModelsTankred Saanum, Peter Dayan, Eric SchulzNeurIPS 2024 · 被引用 14 次
- Evaluating alignment between humans and neural network representations in image-based learning tasksCan Demircan, Tankred Saanum, Leonardo Pettini, Marcel Binz 等NeurIPS 2024 · 被引用 11 次
- Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement LearningYounggyo Seo, Pieter AbbeelNeurIPS 2025 · 被引用 11 次
- Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse RewardJiarui Yang, Bin Zhu, Jingjing Chen, Yu-Gang JiangAAAI 2026 · 被引用 3 次
- Sparse Autoencoders Reveal Temporal Difference Learning in Large Language ModelsCan Demircan, Tankred Saanum, Akshay Kumar Jagadish, Marcel Binz 等ICLR 2025 · 被引用 1 次
它引用的顶会 Paper9
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Improving Sample Efficiency in Model-Free Reinforcement Learning from ImagesDenis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos 等AAAI 2021 · 被引用 506 次
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 被引用 244 次
- Learning Invariant Representations for Reinforcement Learning without ReconstructionAmy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal 等ICLR 2021 · 被引用 77 次
相关 Paper
- Robust Predictable ControlBen Eysenbach, Ruslan Salakhutdinov, Sergey LevineNeurIPS 2021 · 被引用 53 次
- Temporal Predictive Coding For Model-Based Planning In Latent SpaceTung D. Nguyen, Rui Shu, Tuan Pham, Hung Bui 等ICML 2021 · 被引用 65 次
- Efficient Unsupervised Sentence Compression by Fine-tuning Transformers with Reinforcement LearningDemian Gholipour Ghalandari, Chris Hokamp, Georgiana IfrimACL 2022
- Exploration via Planning for Information about the Optimal TrajectoryViraj Mehta, Ian Char, Joseph Abbate, Rory Conlin 等NeurIPS 2022 · 被引用 12 次
- Overcoming Slow Decision Frequencies in Continuous Control: Model-Based Sequence Reinforcement Learning for Model-Free ControlDevdhar Patel, Hava T. SiegelmannICLR 2025
