Generator-Mediated Bandits: Thompson Sampling for GenAI-Powered Adaptive Interventions
Marc Brooks, Gabriel Durham, Kihyuk Hong, Ambuj Tewari
Abstract
Recent advances in generative artificial intelligence (GenAI) models have enabled the generation of personalized content that adapts to up-to-date user context. While personalized decision systems are often modeled using bandit formulations, the integration of GenAI introduces new structure into otherwise classical sequential learning problems. In GenAI-powered interventions, the agent selects a query, but the environment experiences a stochastic response drawn from the generative model. Standard bandit methods do not explicitly account for this structure, where actions influence rewards only through stochastic, observed treatments. We introduce generator-mediated bandit-Thompson sampling (GAMBITTS), a bandit approach designed for this action/treatment split, using mobile health interventions with large language model-generated text as a motivating case study. GAMBITTS explicitly models both the treatment and reward generation processes, using information in the delivered treatment to accelerate policy learning relative to standard methods. We establish regret bounds for GAMBITTS by decomposing sources of uncertainty in treatment and reward, identifying conditions where it achieves stronger guarantees than standard bandit approaches. In simulation studies, GAMBITTS consistently outperforms conventional algorithms by leveraging observed treatments to more accurately estimate expected rewards.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- Personalized HeartSteps: A Reinforcement Learning Algorithm for Optimizing Physical ActivityPeng Liao, Kristjan H. Greenewald, Predrag V. Klasnja, Susan A. MurphyUbiComp 2020 · 163 citations
- Can large language models explore in-context?Akshay Krishnamurthy, Keegan Harris, Dylan J. Foster, Cyril Zhang et al.NeurIPS 2024 · 95 citations
- Efficient Prompt Optimization Through the Lens of Best Arm IdentificationChengshuai Shi, Kun Yang, Zihan Chen, Jundong Li et al.NeurIPS 2024 · 44 citations
- Use Your INSTINCT: INSTruction optimization for LLMs usIng Neural bandits Coupled with TransformersXiaoqiang Lin, Zhaoxuan Wu, Zhongxiang Dai, Wenyang Hu et al.ICML 2024 · 26 citations
Related papers
- The Last JITAI? Exploring Large Language Models for Issuing Just-in-Time Adaptive Interventions: Fostering Physical Activity in a Prospective Cardiac Rehabilitation SettingDavid Haag, Devender Kumar, Sebastian Gruber, Dominik P. Hofer et al.CHI 2025 · 35 citations
- Putting Things into Context: Generative AI-Enabled Context Personalization for Vocabulary Learning Improves Learning MotivationJoanne Leong, Pat Pataranutaporn, Valdemar Danry, Florian Perteneder et al.CHI 2024 · 49 citations
- BFTS: Thompson Sampling with Bayesian Additive Regression TreesRuizhe Deng, Bibhas Chakraborty, Ran Chen, Yan Shuo TanICML 2026 · 1 citation
- Large Language Model-Enhanced Multi-Armed BanditsJiahang Sun, Zhiyong Wang, Runhan Yang, Chenjun Xiao et al.ACL 2026 · 6 citations
- A Comparative Study of Users' Information-Seeking Practices Across Search Engines and Generative AI ChatbotsElsa Lichtenegger, Aleksandra Urman, Aniko HannakSIGIR 2026
