ICML2026
Time-Consistent Robust Multi-Objective Reinforcement Learning via a Bellman–Isaacs Weight-Adversary Recursion
Mingxi Hu, Meiling Yu
摘要
Most multi-objective reinforcement learning (MORL) methods either condition on a fixed preference weight or consider episodic robustness where an adversary selects a single per episode. We study a time-consistent robustness model with reactive preferences: after each transition, an opponent chooses the next weight after observing , and incurs a switching cost based on a Bregman divergence. This yields a Bellman–Isaacs recursion with an inner weight minimization at every backup. We prove the induced operator is a contraction and derive a Bellman-residual certificate that turns approximation error into a uniform bound on robust performance. We develop practical solvers in both tabular and deep settings using Bregman-prox inner updates and a stabilized fixed-point iteration. To evaluate robustness without optimistic critic reuse, we introduce BR-, testing policies against independently trained best-response preference adversaries. Across MO-Gymnasium benchmarks, our approach consistently improves WRR under strong step-wise opponents over preference-conditioned baselines while keeping DRIFT smoothly controllable via .