DMBP: Diffusion model-based predictor for robust offline reinforcement learning against state observation perturbations
Zhihe Yang, Yunjian Xu
Abstract
Offline reinforcement learning (RL), which aims to fully explore offline datasets for training without interaction with environments, has attracted growing recent attention. A major challenge for the real-world application of offline RL stems from the robustness against state observation perturbations, e.g., as a result of sensor errors or adversarial attacks. Unlike online robust RL, agents cannot be adversarially trained in the offline setting. In this work, we propose Diffusion Model-Based Predictor (DMBP) in a new framework that recovers the actual states with conditional diffusion models for state-based RL tasks. To mitigate the error accumulation issue in model-based estimation resulting from the classical training of conventional diffusion models, we propose a non-Markovian training objective to minimize the sum entropy of denoised states in RL trajectory. Experiments on standard benchmark problems demonstrate that DMBP can significantly enhance the robustness of existing offline RL algorithms against different scales of random noises and adversarial attacks on state observations. Further, the proposed framework can effectively deal with incomplete state observations with random combinations of multiple unobserved dimensions in the test. Our implementation is available at https://github.com/zhyang2226/DMBP * Corresponding author Robust RL. Robust RL can be categorized into two taxonomies: training-time and testing-time robustness. Training-time robust RL involves perturbations during the training process, while evaluating the agent in a clean environment (Zhang et al., 2022b; Ye et al., 2023) . Conversely, testing-time robust RL focuses on training the agent with unperturbed datasets or environments and then testing its performance in the presence of disturbances (Yang et al., 2022; Panaganti et al., 2022) . Our work primarily aims at enhancing the testing-time robustness of existing offline RL algorithms. Testing-time robust RL formulations can generally be divided into three categories (Xu et al., 2022) . i) Uncertain observations: In online settings, Zhang et al. (2020) propose a state-adversarial Markov decision process (SA-MDP) framework, which is advanded by Zhang et al. (2021); Sun et al. (2021) that adopt neural networks to simulate worst-case observation attacks for the training of more robust policies. In offline settings, Yang et al. (2022) utilize the conservative smoothing method to make the agent take similar actions when the perturbations on state observation are relatively small. ii) Uncertain actions: Tessler et al. (2019) explore the training of robust policies against two types of action uncertainties, i.e., occasional and constant adversarial perturbations. Tan et al. (2020) utilize adversarial training on actions to enhance the robustness against action perturbations. iii) Uncertain transitions and rewards: The computation of optimal policies against uncertain environment parameters has been explored under the robust Markov Decision Process (MDP) (Xu &
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37cb3df1-445b-46f1-88eb-0c14536ccfc4Cited by top-tier papers9
- Uncertainty-based Offline Variational Bayesian Reinforcement Learning for Robustness under Diverse Data CorruptionsRui Yang, Jie Wang, Guoping Wu, Bin LiNeurIPS 2024 · 11 citations
- Dual-Robust Cross-Domain Offline Reinforcement Learning Against Dynamics ShiftsZhongjian Qiao, Rui Yang, Jiafei Lyu, Xiu Li et al.ICLR 2026 · 7 citations
- ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement LearningZeyuan Liu, Zhihe Yang, Jiawei Xu, Rui Yang et al.NeurIPS 2025 · 3 citations
- Robust Deep Reinforcement Learning against Adversarial Behavior ManipulationShojiro Yamabe, Kazuto Fukuchi, Jun SakumaICLR 2026 · 1 citation
- Forecasting in Offline Reinforcement Learning for Non-stationary EnvironmentsSuzan Ece Ada, Georg Martius, Emre Ugur, Erhan OztopNeurIPS 2025
Builds on23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsAndreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu et al.CVPR 2022 · 1,425 citations
Related papers
- Enhancing Diffusion Policies with Distribution-Matching Generator in Offline Reinforcement LearningXuemin Hu, Shen Li, Yingfen Xu, Bo Tang et al.AAAI 2026 · 1 citation
- Adversarial Diffusion for Robust Reinforcement LearningDaniele Foffano, Alessio Russo, Alexandre ProutièreNeurIPS 2025 · 5 citations
- RORL: Robust Offline Reinforcement Learning via Conservative SmoothingRui Yang, Chenjia Bai, Xiaoteng Ma, Zhaoran Wang et al.NeurIPS 2022 · 118 citations
- MADiff: Offline Multi-agent Learning with Diffusion ModelsZhengbang Zhu, Minghuan Liu, Liyuan Mao, Bingyi Kang et al.NeurIPS 2024 · 116 citations
- Adversarial Model for Offline Reinforcement LearningMohak Bhardwaj, Tengyang Xie, Byron Boots, Nan Jiang et al.NeurIPS 2023 · 44 citations
