Prompt Tuning Decision Transformers with Structured and Scalable Bandits
Finn Rietz, Oleg Smirnov, Sara Karimi, Lele Cao
Abstract
Prompt tuning has emerged as a key technique for adapting large pre-trained Decision Transformers (DTs) in offline Reinforcement Learning (RL), particularly in multi-task and few-shot settings. The Prompting Decision Transformer (PDT) enables task generalization via trajectory prompts sampled uniformly from expert demonstrations -- without accounting for prompt informativeness. In this work, we propose a bandit-based prompt-tuning method that learns to construct optimal trajectory prompts from demonstration data at inference time. We devise a structured bandit architecture operating in the trajectory prompt space, achieving linear rather than combinatorial scaling with prompt size. Additionally, we show that the pre-trained PDT itself can serve as a powerful feature extractor for the bandit, enabling efficient reward modeling across various environments. We theoretically establish regret bounds and demonstrate empirically that our method consistently enhances performance across a wide range of tasks, high-dimensional environments, and out-of-distribution scenarios, outperforming existing baselines in prompt tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ece81cab-cace-48bd-865c-b0d2b379a797Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Neural Contextual Bandits with UCB-based ExplorationDongruo Zhou, Lihong Li, Quanquan GuICML 2020 · 329 citations
- Prompting Decision Transformer for Few-Shot Policy GeneralizationMengdi Xu, Yikang Shen, Shun Zhang, Yuchen Lu et al.ICML 2022 · 194 citations
- Offline Meta-Reinforcement Learning with Advantage WeightingEric Mitchell, Rafael Rafailov, Xue Bin Peng, Sergey Levine et al.ICML 2021 · 122 citations
Related papers
- Decomposed Prompt Decision Transformer for Efficient Unseen Task GeneralizationHongling Zheng, Li Shen, Yong Luo, Tongliang Liu et al.NeurIPS 2024 · 13 citations
- Supervised Pretraining Can Learn In-Context Reinforcement LearningJonathan Lee, Annie Xie, Aldo Pacchiano, Yash Chandak et al.NeurIPS 2023 · 170 citations
- Future-conditioned Unsupervised Pretraining for Decision TransformerZhihui Xie, Zichuan Lin, Deheng Ye, Qiang Fu et al.ICML 2023 · 32 citations
- Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model DisentanglementZhi Wang, Li Zhang, Wenhao Wu, Yuanheng Zhu et al.NeurIPS 2024 · 31 citations
- Pre-Trained Multi-Goal Transformers with Prompt Optimization for Efficient Online AdaptationHaoqi Yuan, Yuhui Fu, Feiyang Xie, Zongqing LuNeurIPS 2024 · 5 citations
