CollabLLM: From Passive Responders to Active Collaborators
Shirley Wu, Michel Galley, Baolin Peng, Hao Cheng, Gavin Li, Yao Dou, Weixin Cai, James Zou, Jure Leskovec, Jianfeng Gao
Abstract
Website: aka.ms/CollabLLM โฃ Multiturn-aware Reward โก Response ๐ Real-world or Simulated User Policy ๐ ๐ฝ ๐ ๐ โข Collaborative Simulation Forward Sampling Reward Computation #1 #2 #3 โ Context state (๐) I need to write about how optimism can improve our well-being. To get us started, what kind of tone are you aiming for? Online generation RL finetuning #1 #2 #3 โฆ โฆ (๐, ๐) Extrinsic Reward e.g., Performance Intrinsic Reward Interactivity Efficiency Figure 1: COLLABLLM Framework: Given a context 1 , the model generates a response 2 to maximize long-term collaboration gains, termed Multiturn-aware Rewards (MR). During training, MRs are estimated via 3 collaborative simulation, which forward-samples conversations with simulated users. Finally, 4 reinforcement fine-tuning is applied using the MRs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6828db0-e10e-48ec-b9b8-b07e184dfc23Cited by top-tier papers17
- Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent CollaborationYijia Shao, Vinay Samuel, Yucheng Jiang, John Yang et al.ICLR 2026 ยท 57 citations
- BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental DesignDeepro Choudhury, Sinead Williamson, Adam Golinski, Ning Miao et al.ICLR 2026 ยท 24 citations
- Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior DataYuxuan Lu, Jing Huang, Yan Han, Bingsheng Yao et al.ACL 2026 ยท 17 citations
- Value of Information: A Framework for Human-Agent CommunicationYijiang River Dong, Tiancheng Hu, Zheng Hui, Caiqi Zhang et al.ACL 2026 ยท 8 citations
- GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI TasksSaelyne Yang, Jaesang Yu, Yi-Hao Peng, Kevin Qinghong Lin et al.CVPR 2026 ยท 5 citations
Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 ยท 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 ยท 10,924 citations
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 ยท 892 citations
- Generating Clarifying Questions for Information RetrievalHamed Zamani, Susan T. Dumais, Nick Craswell, Paul N. Bennett et al.WWW 2020 ยท 238 citations
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue AgentsYang Deng, Wenxuan Zhang, Wai Lam, See-Kiong Ng et al.ICLR 2024 ยท 86 citations
Related papers
- LLM Collaboration with Multi-Agent Reinforcement LearningShuo Liu, Zeyu Liang, Xueguang Lyu, Christopher AmatoAAAI 2026
- MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement LearningChanwoo Park, Seungju Han, Xingzhi Guo, Asuman E. Ozdaglar et al.ACL 2025 ยท 63 citations
- InfoPO: Information-Driven Policy Optimization for User-Centric AgentsFanqi Kong, Jiayi Zhang, Mingyi Deng, Chenglin Wu et al.ICML 2026
- Implicit Turn-Wise Policy Optimization for Proactive User-LLM InteractionHaoyu Wang, Yuxin Chen, Liang Luo, Buyun Zhang et al.ICML 2026 ยท 3 citations
- Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement LearningHao Ma, Tianyi Hu, Zhiqiang Pu, Boyin Liu et al.NeurIPS 2024 ยท 54 citations
