Reflect-then-Correct: Rebalancing Task Optimization for Generalizable Meta-Reinforcement Learning via Distributional Value Error Reduction
Min Wang, Xin Li, Ye He, Mingzhong Wang, Yonggang Zhang
Abstract
Meta-Reinforcement Learning (Meta-RL) faces significant challenges in non-parametric settings, where vastly different return scales across diverse tasks cause severe gradient interference. Existing categorical solutions attempt to normalize these scales but often fail due to rigid discretization and quantization errors. To address this, we propose Reflect-then-Correct (RTC), a framework that models meta-values using Sinkhorn divergence. By treating distributions as adaptive floating particles, RTC achieves a geometry-aware alignment of distinct meta-task structures. However, while Sinkhorn updates harmonize gradients, they introduce statistical bias via sampling estimation. RTC overcomes this issue by "reflecting" on the temporal accumulation of Bellman inconsistencies through a recursive error model and "correcting" the optimization via adaptive importance weights, which prioritize more accurate transitions for meta-value estimation. We provide theoretical guarantees for this reweighting strategy and demonstrate that RTC outperforms existing baselines on the challenging Meta-World ML-10 and ML-45 benchmarks. Source code is available at https://github.com/MinWangcs/RTC
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b7f228d-5737-45b6-a7cd-a732bf49ce74Builds on11
- Faster Wasserstein Distance Estimation with the Sinkhorn DivergenceLénaïc Chizat, Pierre Roussillon, Flavien Léger, François-Xavier Vialard et al.NeurIPS 2020 · 164 citations
- Towards Effective Context for Meta-Reinforcement Learning: an Approach based on Contrastive LearningHaotian Fu, Hongyao Tang, Jianye Hao, Chen Chen et al.AAAI 2021 · 61 citations
- Linear Time Sinkhorn Divergences using Positive FeaturesMeyer Scetbon, Marco CuturiNeurIPS 2020 · 31 citations
- Improving Generalization in Meta-RL with Imaginary Tasks from Latent Dynamics MixtureSuyoung Lee, Sae-Young ChungNeurIPS 2021 · 23 citations
- AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with TransformersJake Grigsby, Justin Sasek, Samyak Parajuli, Daniel Adebi et al.NeurIPS 2024 · 19 citations
Related papers
- Distributional Reinforcement Learning with Regularized Wasserstein LossKe Sun, Yingnan Zhao, Wulong Liu, Bei Jiang et al.NeurIPS 2024 · 2 citations
- Doubly Robust Augmented Transfer for Meta-Reinforcement LearningYuankun Jiang, Nuowen Kan, Chenglin Li, Wenrui Dai et al.NeurIPS 2023 · 3 citations
- Soft Conflict-Resolution Decision Transformer for Offline Multi-Task Reinforcement LearningShudong Wang, Xinfei Wang, Chenhao Zhang, Shanchen Pang et al.AAAI 2026 · 2 citations
- Alleviating Shifted Distribution in Human Preference Alignment through Meta-LearningShihan Dou, Yan Liu, Enyu Zhou, Songyang Gao et al.AAAI 2025 · 2 citations
- Offline Meta Reinforcement Learning with In-Distribution Online AdaptationJianhao Wang, Jin Zhang, Haozhe Jiang, Junyu Zhang et al.ICML 2023 · 16 citations
