Lune

ICML2026Top-tier venue

Time-Consistent Robust Multi-Objective Reinforcement Learning via a Bellman–Isaacs Weight-Adversary Recursion

Mingxi Hu, Meiling Yu

2026Year

Abstract

Most multi-objective reinforcement learning (MORL) methods either condition on a fixed preference weight ww or consider episodic robustness where an adversary selects a single ww per episode. We study a time-consistent robustness model with reactive preferences: after each transition, an opponent chooses the next weight wt+1w_{t+1} after observing st+1s_{t+1}, and incurs a switching cost λDΦ(wt+1∣wt)\lambda D_\Phi(w_{t+1}\mid w_t) based on a Bregman divergence. This yields a Bellman–Isaacs recursion with an inner weight minimization at every backup. We prove the induced operator is a contraction and derive a Bellman-residual certificate that turns approximation error into a uniform bound on robust performance. We develop practical solvers in both tabular and deep settings using Bregman-prox inner updates and a stabilized fixed-point iteration. To evaluate robustness without optimistic critic reuse, we introduce BR-KK, testing policies against KK independently trained best-response preference adversaries. Across MO-Gymnasium benchmarks, our approach consistently improves WRR under strong step-wise opponents over preference-conditioned baselines while keeping DRIFT smoothly controllable via λ\lambda.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 4a8404b5-e471-4848-977d-8c0b85e1d130

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines