Reward-Preserving Counterfactual State Editing for Offline Reinforcement Learning
Siyu Wang, Xiaocong Chen, Mingming Gong, Yong Li, Quan Sheng, Lina Yao
Abstract
Transformer sequence models such as Decision Transformer can learn strong offline policies from logged trajectories, but they often suffer from causal confusion: reliance on spurious correlations that predict reward in the data but do not reflect the true causal mechanisms of the environment. We propose CSET (Counterfactual State Editing Transformer), which improves robustness in strictly offline reinforcement learning without learning environment transition dynamics. On the data side, CSET fits a causal reward model as a conditional variational autoencoder and uses a counterfactual state generator to propose minimally edited observations whose predicted reward matches the factual reward, under a normalized move-band constraint and an acceptance gate that enforce plausibility and reward consistency; augmentation replaces only the observation token to avoid synthetic successor transitions. On the model side, CSET uses a causally structured hybrid transformer: modality-specific convolutional encoders process return-to-go, state, and action streams, and a final attention block is softly supervised so action prediction focuses on its direct causal parents. Experiments on D4RL locomotion, AntMaze, and offline recommendation benchmarks show consistent gains within the DT family, and CSET remains substantially more robust than strong value-based and DT baselines under injected spurious distractors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cda19d91-62bb-4643-aa19-9d3322f179f9Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
Related papers
- Causal Decision Transformer for Recommender Systems via Offline Reinforcement LearningSiyu Wang, Xiaocong Chen, Dietmar Jannach, Lina YaoSIGIR 2023 · 33 citations
- Q-value Regularized Transformer for Offline Reinforcement LearningShengchao Hu, Ziqing Fan, Chaoqin Huang, Li Shen et al.ICML 2024 · 34 citations
- Less is More: an Attention-free Sequence Prediction Modeling for Offline Embodied LearningWei Huang, Jianshu Zhang, Leiyu Wang, Heyue Li et al.NeurIPS 2025
- Addressing Optimism Bias in Sequence Modeling for Reinforcement LearningAdam R. Villaflor, Zhe Huang, Swapnil Pande, John M. Dolan et al.ICML 2022 · 30 citations
- Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RLTaku Yamagata, Ahmed Khalil, Raúl Santos-RodríguezICML 2023 · 121 citations
