Cache-Efficient Posterior Sampling for Reinforcement Learning with LLM-Derived Priors Across Discrete and Continuous Domains
Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
Abstract
Integrating large language models (LLMs) as action proposers in reinforcement learning (RL) boosts performance in text-based environments but incurs high computational costs. We introduce a cache-efficient framework for Bayesian RL with LLM-derived action suggestions, reducing costs while maintaining near-optimal performance. Our approach features a meta-learned adaptive cache, optimized via meta-learning based on policy performance, enabling efficient inference in text-based games (e.g., TextWorld, ALFWorld) and robotic control tasks (e.g., MuJoCo, Meta-World). It achieves a 3.8-4.7× reduction in LLM queries, 4.0-12.0× lower median latencies (85-93ms on consumer hardware), and retains 96-98% of uncached performance. Theoretical KL-divergence bounds ensure reliable cached decisions, validated empirically across tasks with 90.4-95.6% success rates in text environments. For offline RL, our CQL-Prior variant improves performance by 14-29% and reduces training time by 38-40%. Evaluations across eight diverse tasks demonstrate the framework's generalizability and practicality for resource-constrained settings, making LLMguided RL viable for text-based and robotic applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9518f62b-b5e1-4976-af59-8bab4aadc59eCited by top-tier papers1
Ask how each one uses itBuilds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
Related papers
- Efficient Reinforcement Learning with Large Language Model PriorsXue Yan, Yan Song, Xidong Feng, Mengyue Yang et al.ICLR 2025
- Model-Based Imaginative Planning for Embodied AgentsJunru Song, Hengzhe Jin, Yucong Huang, Tingsong Jiang et al.ACL 2026
- Meta-RL Induces Exploration in Language AgentsYulun Jiang, Liangze Jiang, Damien Teney, Michael Moor et al.ICLR 2026 · 20 citations
- Enhancing Decision-Making of Large Language Models via Actor-CriticHeng Dong, Kefei Duan, Chongjie ZhangICML 2025
- Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMsYifan Zhou, Sachin Grover, Mohamed El Mistiri, Kamalesh Kalirathinam et al.NeurIPS 2025 · 3 citations
