QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL
Xing Lei, Jincheng Wang, Xuetao Zhang, Donglin Wang
摘要
Offline goal-conditioned RL (GCRL) learns goal-reaching policies from static datasets, but real-world environments are often partially observable, so the collected trajectories are only partly consistent with the Markov assumption while other segments remain history-dependent. History-aware sequence models such as Decision Transformer (DT) are a natural fit for long-term dependency modeling, yet pure attention is inefficient and brittle when handling local Markovian structure and long-range context simultaneously. Although recent hybrid architectures (e.g., LSDT) introduce local extractors, their fixed-window extraction cannot adapt the effective memory to varying dependency lengths, often truncating long-range context instead of compressing it. Moreover, under sparse rewards, return-to-go (RTG) becomes non-discriminative across sub-trajectories, offering little guidance for stitching goal-reaching behaviors from diverse demonstrations. To address these limitations, we propose QHyer (Q-conditioned Hybrid Attention-Mamba Transformer), which replaces RTG with a Normalizing Flows (NFs) parameterized goal-reaching Q-estimator used directly as conditioning tokens, and a gated Hybrid Attention-Mamba backbone whose selective state-space dynamics enable content-adaptive history compression while attention captures global goal-directed dependencies. Extensive experiments on OGBench and D4RL demonstrate that QHyer achieves state-of-the-art performance on both non-Markovian and Markovian datasets, validating its effectiveness for diverse scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper46
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
相关 Paper
- Long-Short Decision Transformer: Bridging Global and Local Dependencies for Generalized Decision-MakingJincheng Wang, Penny Karanasou, Pengyuan Wei, Elia Gatti 等ICLR 2025
- Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence ModelingSili Huang, Jifeng Hu, Zhejian Yang, Liwei Yang 等NeurIPS 2024 · 被引用 65 次
- Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RLTaku Yamagata, Ahmed Khalil, Raúl Santos-RodríguezICML 2023 · 被引用 121 次
- Less is More: an Attention-free Sequence Prediction Modeling for Offline Embodied LearningWei Huang, Jianshu Zhang, Leiyu Wang, Heyue Li 等NeurIPS 2025
- Decision Mamba: A Multi-Grained State Space Model with Self-Evolution Regularization for Offline RLQi Lv, Xiang Deng, Gongwei Chen, Michael Yu Wang 等NeurIPS 2024 · 被引用 25 次
