Regularity as Intrinsic Reward for Free Play
Cansu Sancaktar, Justus H. Piater, Georg Martius
Abstract
We propose regularity as a novel reward signal for intrinsically-motivated reinforcement learning. Taking inspiration from child development, we postulate that striving for structure and order helps guide exploration towards a subspace of tasks that are not favored by naive uncertainty-based intrinsic rewards. Our generalized formulation of Regularity as Intrinsic Reward (RaIR) allows us to operationalize it within model-based reinforcement learning. In a synthetic environment, we showcase the plethora of structured patterns that can emerge from pursuing this regularity objective. We also demonstrate the strength of our method in a multiobject robotic manipulation environment. We incorporate RaIR into free play and use it to complement the model's epistemic uncertainty as an intrinsic reward. Doing so, we witness the autonomous construction of towers and other regular structures during free play, which leads to a substantial improvement in zero-shot downstream task performance on assembly tasks. Code and videos are available at https://sites.google.com/view/rair-project .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 533e01ed-3e88-4359-aa46-5cfbffe4486aCited by top-tier papers3
- Teaching Models to Teach Themselves: Reasoning at the Edge of LearnabilityShobhita Sundaram, John Quan, Ariel Kwiatkowski, Kartik Ahuja et al.ICML 2026 · 16 citations
- Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement LearningPatrik Reizinger, Bálint Mucsányi, Siyuan Guo, Benjamin Eysenbach et al.ICLR 2026 · 4 citations
- SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World ModelsCansu Sancaktar, Christian Gumbsch, Andrii Zadaianchuk, Pavel Kolev et al.ICML 2025
Builds on8
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- Discovering and Achieving Goals via World ModelsRussell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner et al.NeurIPS 2021 · 177 citations
- Planning Goals for ExplorationEdward S. Hu, Richard Chang, Oleh Rybkin, Dinesh JayaramanICLR 2023 · 152 citations
Related papers
- Curious Exploration via Structured World Models Yields Zero-Shot Object ManipulationCansu Sancaktar, Sebastian Blaes, Georg MartiusNeurIPS 2022 · 43 citations
- Causal Curiosity: RL Agents Discovering Self-supervised Experiments for Causal Representation LearningSumedh A. Sontakke, Arash Mehrjou, Laurent Itti, Bernhard SchölkopfICML 2021 · 73 citations
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 198 citations
- Mega-Reward: Achieving Human-Level Play without Extrinsic RewardsYuhang Song, Jianyi Wang, Thomas Lukasiewicz, Zhenghua Xu et al.AAAI 2020 · 18 citations
- Sequential Generative Exploration Model for Partially Observable Reinforcement LearningHaiyan Yin, Jianda Chen, Sinno Jialin Pan, Sebastian TschiatschekAAAI 2021 · 7 citations
