Test-Time Regret Minimization in Meta Reinforcement Learning
Mirco Mutti, Aviv Tamar
Abstract
Meta reinforcement learning sets a distribution over a set of tasks on which the agent can train at will, then is asked to learn an optimal policy for any test task efficiently. In this paper, we consider a finite set of tasks modeled through Markov decision processes with various dynamics. We assume to have endured a long training phase, from which the set of tasks is perfectly recovered, and we focus on regret minimization against the optimal policy in the unknown test task. Under a separation condition that states the existence of a state-action pair revealing a task against another, Chen et al. (2022) show that regret can be achieved, where are the number of tasks in the set and test episodes, respectively. In our first contribution, we demonstrate that the latter rate is nearly optimal by developing a novel lower bound for test-time regret minimization under separation, showing that a linear dependence with is unavoidable. Then, we present a family of stronger yet reasonable assumptions beyond separation, which we call strong identifiability, enabling algorithms achieving fast rates and sublinear dependence with simultaneously. Our paper provides a new understanding of the statistical barriers of test-time regret minimization and when fast rates can be achieved.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- A Classification View on Meta Learning BanditsMirco Mutti, Jeongyeol Kwon, Shie Mannor, Aviv TamarICML 2025
- Constrained Meta Reinforcement Learning with Provable Test-Time SafetyTingting Ni, Maryam KamgarpourICML 2026
- Test-Time Visual In-Context TuningJiahao Xie, Alessio Tonioni, Nathalie Rauschmayr, Federico Tombari et al.CVPR 2025
Builds on21
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze et al.ICLR 2020 · 315 citations
- Understanding Domain Randomization for Sim-to-real TransferXiaoyu Chen, Jiachen Hu, Chi Jin, Lihong Li et al.ICLR 2022 · 164 citations
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler et al.NeurIPS 2020 · 124 citations
- Fast active learning for pure exploration in reinforcement learningPierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann et al.ICML 2021 · 110 citations
- RL for Latent MDPs: Regret Guarantees and a Lower BoundJeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, Shie MannorNeurIPS 2021 · 91 citations
Related papers
- Model-based Adversarial Meta-Reinforcement LearningZichuan Lin, Garrett Thomas, Guangwen Yang, Tengyu MaNeurIPS 2020 · 58 citations
- Distributionally Adaptive Meta Reinforcement LearningAnurag Ajay, Abhishek Gupta, Dibya Ghosh, Sergey Levine et al.NeurIPS 2022 · 21 citations
- Improving Generalization in Meta-RL with Imaginary Tasks from Latent Dynamics MixtureSuyoung Lee, Sae-Young ChungNeurIPS 2021 · 23 citations
- Meta Reinforcement Learning with Finite Training Tasks - a Density Estimation ApproachZohar Rimon, Aviv Tamar, Gilad AdlerNeurIPS 2022 · 9 citations
- Information-theoretic Task Selection for Meta-Reinforcement LearningRicardo Luna Gutiérrez, Matteo LeonettiNeurIPS 2020 · 24 citations
