Operator World Models for Reinforcement Learning
Pietro Novelli, Marco Pratticò, Massimiliano Pontil, Carlo Ciliberto
Abstract
Policy Mirror Descent (PMD) is a powerful and theoretically sound methodology for sequential decision-making. However, it is not directly applicable to Reinforcement Learning (RL) due to the inaccessibility of explicit action-value functions. We address this challenge by introducing a novel approach based on learning a world model of the environment using conditional mean embeddings. Leveraging tools from operator theory we derive a closed-form expression of the action-value function in terms of the world model via simple matrix operations. Combining these estimators with PMD leads to POWR, a new RL algorithm for which we prove convergence rates to the global optimum. Preliminary experiments in finite and infinite state settings support the effectiveness of our method
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b36e314-4c05-463e-aaae-7d64bc804da5Cited by top-tier papers3
- Self-Supervised Evolution Operator Learning for High-Dimensional Dynamical SystemsGiacomo Turri, Luigi Bonati, Kai Zhu, Massimiliano Pontil et al.ICLR 2026 · 10 citations
- Sequence Modeling with Spectral Mean FlowsJinwoo Kim, Max Beier, Petar Bevanda, Nayun Kim et al.NeurIPS 2025 · 2 citations
- Conditional KRR: Injecting Unpenalized Features into Kernel Methods with Applications to Kernel ThresholdingRustem Takhanov, Zhenisbek AssylbekovICML 2026
Builds on7
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang et al.ICML 2020 · 324 citations
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 201 citations
- Learning Dynamical Systems via Koopman Operator Regression in Reproducing Kernel Hilbert SpacesVladimir Kostic, Pietro Novelli, Andreas Maurer, Carlo Ciliberto et al.NeurIPS 2022 · 109 citations
- Optimal Rates for Regularized Conditional Mean Embedding LearningZhu Li, Dimitri Meunier, Mattes Mollenhauer, Arthur GrettonNeurIPS 2022 · 69 citations
- Sharp Spectral Rates for Koopman Operator LearningVladimir Kostic, Karim Lounici, Pietro Novelli, Massimiliano PontilNeurIPS 2023 · 57 citations
Related papers
- Fast Adaptation to New Environments via Policy-Dynamics Value FunctionsRoberta Raileanu, Maxwell Goldstein, Arthur Szlam, Rob FergusICML 2020 · 27 citations
- PWM: Policy Learning with Multi-Task World ModelsIgnat Georgiev, Varun Giridhar, Nicklas Hansen, Animesh GargICLR 2025
- Breaking the Computational Barrier: Provably Efficient Actor–Critic for Low-Rank MDPsRuiquan Huang, Donghao Li, Yingbin LIANG, Jing YangICML 2026
- Rank-One Modified Value IterationArman Sharifi Kolarijani, Tolga Ok, Peyman Mohajerin Esfahani, Mohamad Amin Sharifi KolarijaniICML 2025
- Inverse Reinforcement Learning with the Average Reward CriterionFeiyang Wu, Jingyang Ke, Anqi WuNeurIPS 2023 · 16 citations
