Fine-grained List-wise Alignment for Generative Medication Recommendation
Chenxiao Fan, Chongming Gao, Wentao Shi, Yaxin Gong, Zihao Zhao, Fuli Feng
Abstract
Accurate and safe medication recommendations are critical for effective clinical decision-making, especially in multimorbidity cases. However, existing systems rely on point-wise prediction paradigms that overlook synergistic drug effects and potential adverse drug-drug interactions (DDIs). We propose FLAME, a fine-grained list-wise alignment framework for large language models (LLMs), enabling drug-by-drug generation of drug lists. FLAME formulates recommendation as a sequential decision process, where each step adds or removes a single drug. To provide fine-grained learning signals, we devise step-wise Group Relative Policy Optimization (GRPO) with potential-based reward shaping, which explicitly models DDIs and optimizes the contribution of each drug to the overall prescription. Furthermore, FLAME enhances patient modeling by integrating structured clinical knowledge and collaborative information into the representation space of LLMs. Experiments on benchmark datasets demonstrate that FLAME achieves state-of-the-art performance, delivering superior accuracy, controllable safety-accuracy trade-offs, and strong generalization across diverse clinical scenarios. Our code is available at https://github.com/cxfann/Flame.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1dc20f6-a6a5-4be6-8ac9-ed5c2ff75afeCited by top-tier papers3
- Don't Start Over: A Cost-Effective Framework for Migrating Personalized Prompts Between LLMsZiyi Zhao, Chongming Gao, Yang Zhang, Haoyan Liu et al.AAAI 2026 · 6 citations
- Uncertainty-aware Generative RecommendationChenxiao Fan, Chongming Gao, Yaxin Gong, Haoyan Liu et al.KDD 2026 · 2 citations
- Awaken the Giant: Activating LLMs via Deep Model Guidance for Boundary-aware Medication RecommendationHang Lv, Zixuan Guo, Yanchao Tan, Wanzi Shao et al.KDD 2026
Builds on9
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Conditional Generation Net for Medication RecommendationRui Wu, Zhaopeng Qiu, Jiacheng Jiang, Guilin Qi et al.WWW 2022 · 135 citations
- LLaRA: Large Language-Recommendation AssistantJiayi Liao, Sihang Li, Zhengyi Yang, Jiancan Wu et al.SIGIR 2024 · 120 citations
- Mole-BERT: Rethinking Pre-training Graph Neural Networks for MoleculesJun Xia, Chengshuai Zhao, Bozhen Hu, Zhangyang Gao et al.ICLR 2023 · 119 citations
Related papers
- SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy OptimizationXuankun Rong, Wenke Huang, Tingfeng Wang, Daiguo Zhou et al.CVPR 2026 · 13 citations
- Factorized Latent Reasoning for LLM-based RecommendationTianqi Gao, Chengkai Huang, Zihan Wang, Cao Liu et al.SIGIR 2026
- Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM AlignmentRuoxi Cheng, Haoxuan Ma, Weixin Wang, Ranjie Duan et al.ICLR 2026 · 23 citations
- Align³GR: Unified Multi-Level Alignment for LLM-based Generative RecommendationWencai Ye, Mingjie Sun, Shuhang Chen, Wenjin Wu et al.AAAI 2026 · 2 citations
- RES-MR: Risk-Aware Reasoning for Explainable and Safe Medication RecommendationCong Wang, Jin Li, Shoujin Wang, Yishuo Li et al.SIGIR 2026 · 1 citation
