A Decision-Language Model (DLM) for Dynamic Restless Multi-Armed Bandit Tasks in Public Health
Nikhil Behari, Edwin Zhang, Yunfan Zhao, Aparna Taneja, Dheeraj Nagaraj, Milind Tambe
摘要
Restless multi-armed bandits (RMAB) have demonstrated success in optimizing resource allocation for large beneficiary populations in public health settings. Unfortunately, RMAB models lack flexibility to adapt to evolving public health policy priorities. Concurrently, Large Language Models (LLMs) have emerged as adept automated planners across domains of robotic control and navigation. In this paper, we propose a Decision Language Model (DLM) for RMABs, enabling dynamic fine-tuning of RMAB policies in public health settings using human-language commands. We propose using LLMs as automated planners to (1) interpret human policy preference prompts, (2) propose reward functions as code for a multi-agent RMAB environment, and (3) iterate on the generated reward functions using feedback from grounded RMAB simulations. We illustrate the application of DLM in collaboration with ARMMAN, an India-based non-profit promoting preventative care for pregnant mothers, that currently relies on RMAB policies to optimally allocate health worker calls to low-resource populations. We conduct a technology demonstration in simulation using the Gemini Pro model, showing DLM can dynamically shape policy outcomes using only human prompts as input.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- The Bandit Whisperer: Communication Learning for Restless BanditsYunfan Zhao, Tonghan Wang, Dheeraj Mysore Nagaraj, Aparna Taneja 等AAAI 2025 · 被引用 6 次
- VORTEX: Aligning Task Utility and Human Preferences Through LLM-Guided Reward ShapingGuojun Xiong, Milind TambeAAAI 2026 · 被引用 2 次
- Efficient Evolutionary Search Over Chemical Space with Large Language ModelsHaorui Wang, Marta Skreta, Cher Tian Ser, Wenhao Gao 等ICLR 2025
- KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent SystemsJusheng Zhang, Zimeng Huang, Yijia Fan, Ningyuan Liu 等ICML 2025
它引用的顶会 Paper13
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Eureka: Human-Level Reward Design via Coding Large Language ModelsYecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang 等ICLR 2024 · 被引用 582 次
- Pre-Trained Language Models for Interactive Decision-MakingShuang Li, Xavier Puig, Chris Paxton, Yilun Du 等NeurIPS 2022 · 被引用 341 次
- Guiding Pretraining in Reinforcement Learning with Large Language ModelsYuqing Du, Olivia Watkins, Zihan Wang, Cédric Colas 等ICML 2023 · 被引用 257 次
相关 Paper
- Efficient Sequential Decision Making with Large Language ModelsDingyang Chen, Qi Zhang, Yinglun ZhuEMNLP 2024 · 被引用 3 次
- DLM: Unified Decision Language Models for Offline Multi-Agent Sequential Decision MakingZhuohui Zhang, Bin Cheng, Bin HeICML 2026
- Large Language Model-Enhanced Multi-Armed BanditsJiahang Sun, Zhiyong Wang, Runhan Yang, Chenjun Xiao 等ACL 2026 · 被引用 6 次
- EVOLvE: Evaluating and Optimizing LLMs For In-Context ExplorationAllen Nie, Yi Su, Bo Chang, Jonathan Lee 等ICML 2025
- On the Modeling Capabilities of Large Language Models for Sequential Decision MakingMartin Klissarov, R. Devon Hjelm, Alexander T. Toshev, Bogdan MazoureICLR 2025
