Policy-Guided Causal State Representation for Offline Reinforcement Learning Recommendation
Siyu Wang, Xiaocong Chen, Lina Yao
Abstract
In offline reinforcement learning-based recommender systems (RLRS), learning effective state representations is crucial for capturing user preferences that directly impact long-term rewards. However, raw state representations often contain high-dimensional, noisy information and components that are not causally relevant to the reward. Additionally, missing transitions in offline data make it challenging to accurately identify features that are most relevant to user satisfaction. To address these challenges, we propose Policy-Guided Causal Representation (PGCR), a novel two-stage framework for causal feature selection and state representation learning in offline RLRS. In the first stage, we learn a causal feature selection policy that generates modified states by isolating and retaining only the causally relevant components (CRCs) while altering irrelevant components. This policy is guided by a reward function based on the Wasserstein distance, which measures the causal effect of state components on the reward and encourages the preservation of CRCs that directly influence user interests. In the second stage, we train an encoder to learn compact state representations by minimizing the mean squared error (MSE) loss between the latent representations of the original and modified states, ensuring that the representations focus on CRCs. We provide a theoretical analysis proving the identifiability of causal effects from interventions, validating the ability of PGCR to isolate critical state components for decision-making. Extensive experiments demonstrate that PGCR significantly improves recommendation performance, confirming its effectiveness for offline RL-based recommender systems. CCS Concepts • Information systems → Recommender systems; • Computing methodologies → Reinforcement learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c60b57ce-8aec-4d56-b2d4-dd13acd0b399Builds on8
- Causal Intervention for Leveraging Popularity Bias in RecommendationYang Zhang, Fuli Feng, Xiangnan He, Tianxin Wei et al.SIGIR 2021 · 431 citations
- Learning Invariant Representations for Reinforcement Learning without ReconstructionAmy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal et al.ICLR 2021 · 77 citations
- Removing Hidden Confounding in Recommendation: A Unified Multi-Task Learning ApproachHaoxuan Li, Kunhan Wu, Chunyuan Zheng, Yanghao Xiao et al.NeurIPS 2023 · 68 citations
- Alleviating Matthew Effect of Offline Reinforcement Learning in Interactive RecommendationChongming Gao, Kexin Huang, Jiawei Chen, Yuan Zhang et al.SIGIR 2023 · 65 citations
- Action-Sufficient State Representation Learning for Control with Structural ConstraintsBiwei Huang, Chaochao Lu, Liu Leqi, José Miguel Hernández-Lobato et al.ICML 2022 · 41 citations
Related papers
- Causal Decision Transformer for Recommender Systems via Offline Reinforcement LearningSiyu Wang, Xiaocong Chen, Dietmar Jannach, Lina YaoSIGIR 2023 · 33 citations
- What are the Statistical Limits of Offline RL with Linear Function Approximation?Ruosong Wang, Dean P. Foster, Sham M. KakadeICLR 2021 · 172 citations
- Learning Pseudometric-based Action Representations for Offline Reinforcement LearningPengjie Gu, Mengchen Zhao, Chen Chen, Dong Li et al.ICML 2022 · 17 citations
- BECAUSE: Bilinear Causal Representation for Generalizable Offline Model-based Reinforcement LearningHaohong Lin, Wenhao Ding, Jian Chen, Laixi Shi et al.NeurIPS 2024 · 5 citations
- Contrastive State Augmentations for Reinforcement Learning-Based Recommender SystemsZhaochun Ren, Na Huang, Yidan Wang, Pengjie Ren et al.SIGIR 2023 · 20 citations
