Personalized Reinforcement Learning with a Budget of Policies
Dmitry Ivanov, Omer Ben-Porat
Abstract
Personalization in machine learning (ML) tailors models' decisions to the individual characteristics of users. While this approach has seen success in areas like recommender systems, its expansion into high-stakes fields such as healthcare and autonomous driving is hindered by the extensive regulatory approval processes involved. To address this challenge, we propose a novel framework termed represented Markov Decision Processes (r-MDPs) that is designed to balance the need for personalization with the regulatory constraints. In an r-MDP, we cater to a diverse user population, each with unique preferences, through interaction with a small set of representative policies. Our objective is twofold: efficiently match each user to an appropriate representative policy and simultaneously optimize these policies to maximize overall social welfare. We develop two deep reinforcement learning algorithms that efficiently solve r-MDPs. These algorithms draw inspiration from the principles of classic K-means clustering and are underpinned by robust theoretical foundations. Our empirical investigations, conducted across a variety of simulated environments, showcase the algorithms' ability to facilitate meaningful personalization even under constrained policy budgets. Furthermore, they demonstrate scalability, efficiently adapting to larger policy budgets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0d8ccd1-ec6e-4a93-83ff-808a8e00a16cCited by top-tier papers2
- Learning Policy Committees for Effective Personalization in MDPs with Diverse TasksLuise Ge, Michael Lanier, Anindya Sarkar, Bengisu Guresti et al.ICML 2025
- Structure Detection for Contextual Reinforcement LearningTianyue Zhou, Jung-Hoon Cho, Cathy WuAAAI 2026
Builds on1
Related papers
- Near-Optimal Sample Complexity for Online Constrained MDPsChang Liu, Yunfan Li, Lin F. YangNeurIPS 2025 · 1 citation
- Learning Fair Policies in Multi-Objective (Deep) Reinforcement Learning with Average and Discounted RewardsUmer Siddique, Paul Weng, Matthieu ZimmerICML 2020 · 1 citation
- Local Differential Privacy for Regret Minimization in Reinforcement LearningEvrard Garcelon, Vianney Perchet, Ciara Pike-Burke, Matteo PirottaNeurIPS 2021 · 47 citations
- Policy Optimization with Advantage Regularization for Long-Term Fairness in Decision SystemsEric Yang Yu, Zhizhen Qin, Min Kyung Lee, Sicun GaoNeurIPS 2022 · 18 citations
- Policy Gradient With Serial Markov Chain ReasoningEdoardo Cetin, Oya ÇeliktutanNeurIPS 2022 · 4 citations
