Operator World Models for Reinforcement Learning
Pietro Novelli, Marco Pratticò, Massimiliano Pontil, Carlo Ciliberto
摘要
Policy Mirror Descent (PMD) is a powerful and theoretically sound methodology for sequential decision-making. However, it is not directly applicable to Reinforcement Learning (RL) due to the inaccessibility of explicit action-value functions. We address this challenge by introducing a novel approach based on learning a world model of the environment using conditional mean embeddings. Leveraging tools from operator theory we derive a closed-form expression of the action-value function in terms of the world model via simple matrix operations. Combining these estimators with PMD leads to POWR, a new RL algorithm for which we prove convergence rates to the global optimum. Preliminary experiments in finite and infinite state settings support the effectiveness of our method
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Self-Supervised Evolution Operator Learning for High-Dimensional Dynamical SystemsGiacomo Turri, Luigi Bonati, Kai Zhu, Massimiliano Pontil 等ICLR 2026 · 被引用 10 次
- Sequence Modeling with Spectral Mean FlowsJinwoo Kim, Max Beier, Petar Bevanda, Nayun Kim 等NeurIPS 2025 · 被引用 2 次
- Conditional KRR: Injecting Unpenalized Features into Kernel Methods with Applications to Kernel ThresholdingRustem Takhanov, Zhenisbek AssylbekovICML 2026
它引用的顶会 Paper7
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang 等ICML 2020 · 被引用 324 次
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 被引用 201 次
- Learning Dynamical Systems via Koopman Operator Regression in Reproducing Kernel Hilbert SpacesVladimir Kostic, Pietro Novelli, Andreas Maurer, Carlo Ciliberto 等NeurIPS 2022 · 被引用 109 次
- Optimal Rates for Regularized Conditional Mean Embedding LearningZhu Li, Dimitri Meunier, Mattes Mollenhauer, Arthur GrettonNeurIPS 2022 · 被引用 69 次
- Sharp Spectral Rates for Koopman Operator LearningVladimir Kostic, Karim Lounici, Pietro Novelli, Massimiliano PontilNeurIPS 2023 · 被引用 57 次
相关 Paper
- Fast Adaptation to New Environments via Policy-Dynamics Value FunctionsRoberta Raileanu, Maxwell Goldstein, Arthur Szlam, Rob FergusICML 2020 · 被引用 27 次
- PWM: Policy Learning with Multi-Task World ModelsIgnat Georgiev, Varun Giridhar, Nicklas Hansen, Animesh GargICLR 2025
- Breaking the Computational Barrier: Provably Efficient Actor–Critic for Low-Rank MDPsRuiquan Huang, Donghao Li, Yingbin LIANG, Jing YangICML 2026
- Rank-One Modified Value IterationArman Sharifi Kolarijani, Tolga Ok, Peyman Mohajerin Esfahani, Mohamad Amin Sharifi KolarijaniICML 2025
- Inverse Reinforcement Learning with the Average Reward CriterionFeiyang Wu, Jingyang Ke, Anqi WuNeurIPS 2023 · 被引用 16 次
