Distributional Reinforcement Learning with Regularized Wasserstein Loss
Ke Sun, Yingnan Zhao, Wulong Liu, Bei Jiang, Linglong Kong
Abstract
The empirical success of distributional reinforcement learning (RL) highly relies on the choice of distribution divergence equipped with an appropriate distribution representation. In this paper, we propose Sinkhorn distributional RL (SinkhornDRL), which leverages Sinkhorn divergence—a regularized Wasserstein loss—to minimize the difference between current and target Bellman return distributions. Theoretically, we prove the contraction properties of SinkhornDRL, aligning with the interpolation nature of Sinkhorn divergence between Wasserstein distance and Maximum Mean Discrepancy (MMD). The introduced SinkhornDRL enriches the family of distributional RL algorithms, contributing to interpreting the algorithm behaviors compared with existing approaches by our investigation into their relationships. Empirically, we show that SinkhornDRL consistently outperforms or matches existing algorithms on the Atari games suite and particularly stands out in the multi-dimensional reward setting. .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 088ab7ad-c6a2-4c3d-9d4c-ca2ea999b0f6Cited by top-tier papers5
- Q#: Provably Optimal Distributional RL for LLM Post-TrainingJin Peng Zhou, Kaiwen Wang, Jonathan D. Chang, Zhaolin Gao et al.NeurIPS 2025 · 18 citations
- A Finite Sample Analysis of Distributional TD Learning with Linear Function ApproximationYang Peng, Kaicheng Jin, Liangyu Zhang, Zhihua ZhangNeurIPS 2025 · 6 citations
- Reflect-then-Correct: Rebalancing Task Optimization for Generalizable Meta-Reinforcement Learning via Distributional Value Error ReductionMin Wang, Xin Li, Ye He, Mingzhong Wang et al.ICML 2026
- A Principled Path to Fitted Distributional EvaluationSungee Hong, Jiayi Wang, Zhengling Qi, Raymond K. W. WongNeurIPS 2025
- Distributional value gradients for stochastic environmentsBaptiste Debes, Tinne TuytelaarsICLR 2026
Builds on18
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Unbalanced minibatch Optimal Transport; applications to Domain AdaptationKilian Fatras, Thibault Séjourné, Rémi Flamary, Nicolas CourtyICML 2021 · 183 citations
- Conservative Offline Distributional Reinforcement LearningYecheng Jason Ma, Dinesh Jayaraman, Osbert BastaniNeurIPS 2021 · 118 citations
- Don't Generate Me: Training Differentially Private Generative Models with Sinkhorn DivergenceTianshi Cao, Alex Bie, Arash Vahdat, Sanja Fidler et al.NeurIPS 2021 · 88 citations
- Non-Crossing Quantile Regression for Distributional Reinforcement LearningFan Zhou, Jianing Wang, Xingdong FengNeurIPS 2020 · 63 citations
Related papers
- Multivariate Distributional Reinforcement Learning Using Sliced DivergencesBaptiste Debes, Tinne TuytelaarsICML 2026
- Distributional Reinforcement Learning via Moment MatchingThanh Nguyen-Tang, Sunil Gupta, Svetha VenkateshAAAI 2021 · 44 citations
- Distributional Reinforcement Learning with Monotonic SplinesYudong Luo, Guiliang Liu, Haonan Duan, Oliver Schulte et al.ICLR 2022 · 18 citations
- Bayesian Distributional Policy GradientsLuchen Li, A. Aldo FaisalAAAI 2021 · 11 citations
- Diverse Projection Ensembles for Distributional Reinforcement LearningMoritz Akiya Zanger, Wendelin Boehmer, Matthijs T. J. SpaanICLR 2024 · 9 citations
